Lars van der Laan, Mark Van Der Laan
arXiv 23 Apr 2026 · Statistics — Machine Learning
arXiv:2604.21260 · PDF · DOI · OpenAlex · Extracted main text
We study semisupervised mean estimation with a small labeled sample, a large unlabeled sample, and a black-box prediction model whose output may be miscalibrated. A standard approach in this setting is augmented inverse-probability weighting (AIPW) [Robins et al., 1994], which protects against prediction-model misspecification but can be inefficient when the prediction score is poorly aligned with the outcome scale. We introduce Calibrated Prediction-Powered Inference, which post-hoc calibrates the prediction score on the labeled sample before using it for semisupervised estimation. This simple step requires no retraining and can improve the original score both as a predictor of the outcome and as a regression adjustment for semisupervised inference. We study both linear and isotonic calibration. For isotonic calibration, we establish first-order optimality guarantees: isotonic post-processing can improve predictive accuracy and estimator efficiency relative to the original score and simpler post-processing rules, while no further post-processing of the fitted isotonic score yields additional first-order gains. For linear calibration, we show first-order equivalence to PPI++. We also clarify the relationship among existing estimators, showing that the original PPI estimator is a special case of AIPW and can be inefficient when the prediction model is accurate, while PPI++ is AIPW with empirical efficiency maximization [Rubin et al., 2008]. In simulations and real-data experiments, our calibrated estimators often outperform PPI and are competitive with, or outperform, AIPW and PPI++. We provide an accompanying Python package, ppi_aipw, at https://larsvanderlaan.github.io/ppi-aipw/.
appendix boundary found by appendix_command · 45% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Chernozhukov, Victor and Chetverikov, Denis and Demirer, Mert and Du… (2018) Double/debiased machine learning for treatment and structural parameters | 1.000 | 7 | 3 | 100% |
| 2 | Moore, Kelly L and van der Laan, Mark J (2009) Covariate adjustment in randomized trials with binary outcomes: targeted maximum likelihood estimation | 1.000 | 6 | 3 | 100% |
| 3 | Angelopoulos, Anastasios N. and Bates, Stephen and Fannjiang, Clara… (2023) Prediction-powered inference | 1.000 | 6 | 3 | 100% |
| 4 | Angelopoulos, Anastasios N and Duchi, John C and Zrnic, Tijana (2023) Ppi++: Efficient prediction-powered inference | 1.000 | 5 | 4 | 100% |
| 5 | Ji, Wenlong and Lei, Lihua and Zrnic, Tijana (2025) Predictions as surrogates: Revisiting surrogate outcomes in the age of ai | 1.000 | 5 | 3 | 100% |
| 6 | Rubin, Daniel B and van der Laan, Mark J (2008) Empirical efficiency maximization: improved locally efficient covariate adjustment in randomized experiments and survival analysis | 0.961 | 9 | 6 | 89% |
| 7 | Robins, James M and Rotnitzky, Andrea and Zhao, Lue Ping (1994) Estimation of regression coefficients when some regressors are not always observed | 0.950 | 14 | 5 | 86% |
| 8 | Højbjerre-Frandsen, Emilie and van der Laan, Mark J and Schuler, Ale… (2025) Powering rcts for marginal effects with glms using prognostic score adjustment | 0.950 | 7 | 5 | 86% |
| 9 | Schuler, Alejandro and Walsh, David and Hall, Diana and Walsh, Jon a… (2022) Increasing the efficiency of randomized trial estimates via linear adjustment for a prognostic score | 0.950 | 7 | 5 | 86% |
| 10 | Hansen, Ben B (2008) The prognostic analogue of the propensity score | 0.941 | 6 | 4 | 83% |
Showing the top 10 of 185 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | AI-Assisted Variance Reduction in Randomized Experiments | 0.737 | 3 | 2 |