Stephen Salerno, Zhenke Wu, Tyler McCormick
arXiv 11 Mar 2026 · Statistics — Machine Learning
arXiv:2603.11368 · PDF · DOI · OpenAlex · Extracted main text
When outcome data are expensive or onerous to collect, scientists increasingly substitute predictions from machine learning and AI models for unlabeled cases, a process which has consequences for downstream statistical inference. While recent methods provide valid uncertainty quantification under independent sampling, real-world applications involve missing at random (MAR) labeling and spatial dependence. For inference in this setting, we propose a doubly robust estimator with cross-fit nuisances. We show that cross-fitting induces fold-level correlation that distorts spatial variance estimators, producing unstable or overly conservative confidence intervals. To address this, we propose a jackknife spatial heteroscedasticity and autocorrelation consistent (HAC) variance correction that separates spatial dependence from fold-induced noise. Under standard identification and dependence conditions, the resulting intervals are asymptotically valid. Simulations and benchmark datasets show substantial improvement in finite-sample calibration, particularly under MAR labeling and clustered sampling.
appendix boundary found by appendix_command · 63% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Zrnic, Tijana and Candès, Emmanuel J (2024) Cross-prediction-powered inference | 1.000 | 6 | 4 | 100% |
| 2 | Conley, Timothy G (1999) GMM estimation with cross sectional dependence | 1.000 | 5 | 3 | 100% |
| 3 | Chandrasekhar, Arun G and Jackson, Matthew O and McCormick, Tyler H… (2023) General covariance-based conditions for central limit theorems with dependent triangular arrays self | 0.928 | 4 | 4 | 100% |
| 4 | Bester, C Alan and Conley, Timothy G and Hansen, Christian B (2011) Inference with dependent data using cluster covariance estimators | 0.928 | 4 | 3 | 100% |
| 5 | Chernozhukov, Victor and Chetverikov, Denis and Demirer, Mert and Du… (2018) Double/debiased machine learning for treatment and structural parameters | 0.843 | 4 | 3 | 75% |
| 6 | Bang, Heejung and Robins, James M (2005) Doubly robust estimation in missing data and causal inference models | 0.737 | 3 | 3 | 67% |
| 7 | Angelopoulos, Anastasios N and Duchi, John C and Zrnic, Tijana (2023) PPI++: Efficient prediction-powered inference | 0.644 | 2 | 2 | 100% |
| 8 | Bullock, Eric L and Woodcock, Curtis E and Souza Jr, Carlos and Olof… (2020) Satellite-based estimates reveal widespread forest degradation in the Amazon | 0.644 | 2 | 2 | 100% |
| 9 | Cameron, A Colin and Gelbach, Jonah B and Miller, Douglas L (2011) Robust inference with multiway clustering | 0.644 | 2 | 2 | 100% |
| 10 | Jenish, Nazgul and Prucha, Ingmar R (2009) Central limit theorems and uniform laws of large numbers for arrays of random fields | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 35 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | When Surveys Become Conversations: Adaptive Matrix Validation for AI-Assisted Interviews | 0.000 | 1 | 1 |