EconBase
← All papers

Calibeating Prediction-Powered Inference

Lars van der Laan, Mark Van Der Laan

arXiv 23 Apr 2026 · Statistics — Machine Learning

arXiv:2604.21260 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We study semisupervised mean estimation with a small labeled sample, a large unlabeled sample, and a black-box prediction model whose output may be miscalibrated. A standard approach in this setting is augmented inverse-probability weighting (AIPW) [Robins et al., 1994], which protects against prediction-model misspecification but can be inefficient when the prediction score is poorly aligned with the outcome scale. We introduce Calibrated Prediction-Powered Inference, which post-hoc calibrates the prediction score on the labeled sample before using it for semisupervised estimation. This simple step requires no retraining and can improve the original score both as a predictor of the outcome and as a regression adjustment for semisupervised inference. We study both linear and isotonic calibration. For isotonic calibration, we establish first-order optimality guarantees: isotonic post-processing can improve predictive accuracy and estimator efficiency relative to the original score and simpler post-processing rules, while no further post-processing of the fitted isotonic score yields additional first-order gains. For linear calibration, we show first-order equivalence to PPI++. We also clarify the relationship among existing estimators, showing that the original PPI estimator is a special case of AIPW and can be inefficient when the prediction model is accurate, while PPI++ is AIPW with empirical efficiency maximization [Rubin et al., 2008]. In simulations and real-data experiments, our calibrated estimators often outperform PPI and are competitive with, or outperform, AIPW and PPI++. We provide an accompanying Python package, ppi_aipw, at https://larsvanderlaan.github.io/ppi-aipw/.

Citation extraction

185
references
395
in-text mentions
185
distinct cited
2
self-citations
13,422
main-text words

appendix boundary found by appendix_command · 45% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Chernozhukov, Victor and Chetverikov, Denis and Demirer, Mert and Du… (2018) Double/debiased machine learning for treatment and structural parameters1.00073100%
2Moore, Kelly L and van der Laan, Mark J (2009) Covariate adjustment in randomized trials with binary outcomes: targeted maximum likelihood estimation1.00063100%
3Angelopoulos, Anastasios N. and Bates, Stephen and Fannjiang, Clara… (2023) Prediction-powered inference1.00063100%
4Angelopoulos, Anastasios N and Duchi, John C and Zrnic, Tijana (2023) Ppi++: Efficient prediction-powered inference1.00054100%
5Ji, Wenlong and Lei, Lihua and Zrnic, Tijana (2025) Predictions as surrogates: Revisiting surrogate outcomes in the age of ai1.00053100%
6Rubin, Daniel B and van der Laan, Mark J (2008) Empirical efficiency maximization: improved locally efficient covariate adjustment in randomized experiments and survival analysis0.9619689%
7Robins, James M and Rotnitzky, Andrea and Zhao, Lue Ping (1994) Estimation of regression coefficients when some regressors are not always observed0.95014586%
8Højbjerre-Frandsen, Emilie and van der Laan, Mark J and Schuler, Ale… (2025) Powering rcts for marginal effects with glms using prognostic score adjustment0.9507586%
9Schuler, Alejandro and Walsh, David and Hall, Diana and Walsh, Jon a… (2022) Increasing the efficiency of randomized trial estimates via linear adjustment for a prognostic score0.9507586%
10Hansen, Ben B (2008) The prognostic analogue of the propensity score0.9416483%

Showing the top 10 of 185 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1AI-Assisted Variance Reduction in Randomized Experiments0.73732