Lars van der Laan, Nathan Kallus
arXiv 29 Dec 2025 · Statistics — Machine Learning
arXiv:2512.23694 · PDF · DOI · OpenAlex · Extracted main text
We introduce Iterated Bellman Calibration, a simple, model-agnostic, post-hoc procedure for calibrating off-policy value predictions in infinite-horizon Markov decision processes. Bellman calibration requires that states with similar predicted long-term returns exhibit one-step returns consistent with the Bellman equation under the target policy. We adapt classical histogram and isotonic calibration to the dynamic, counterfactual setting by repeatedly regressing fitted Bellman targets onto a model's predictions, using a doubly robust pseudo-outcome to handle off-policy data. This yields a one-dimensional fitted value iteration scheme that can be applied to any value estimator. Our analysis provides finite-sample guarantees for both calibration and prediction under weak assumptions, and critically, without requiring Bellman completeness or realizability.
appendix boundary found by appendix_command · 38% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Rémi Munos and Csaba Szepesvári (2008) Finite-time bounds for fitted value iteration | 1.000 | 5 | 3 | 100% |
| 2 | Alexandru Niculescu-Mizil and Rich Caruana (2005) Predicting good probabilities with supervised learning | 0.928 | 4 | 3 | 100% |
| 3 | Bianca Zadrozny and Charles Elkan (2002) Transforming classifier scores into accurate multiclass probability estimates | 0.928 | 4 | 3 | 100% |
| 4 | Lars Van Der Laan, Ernesto Ulloa-Pérez, Marco Carone, and Alex Luedtke (2023) Causal isotonic calibration for heterogeneous treatment effects | 0.920 | 9 | 6 | 78% |
| 5 | Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger (2017) On calibration of modern neural networks | 0.843 | 4 | 3 | 75% |
| 6 | Jinglin Chen and Nan Jiang (2019) Information-theoretic considerations in batch reinforcement learning | 0.843 | 3 | 3 | 100% |
| 7 | Chirag Gupta, Aleksandr Podkopaev, and Aaditya Ramdas (2020) Distribution-free binary classification: prediction sets, confidence intervals and calibration | 0.843 | 3 | 3 | 100% |
| 8 | Tengyang Xie and Nan Jiang (2021) Batch value-function approximation with only realizability | 0.811 | 4 | 2 | 100% |
| 9 | Bianca Zadrozny and Charles Elkan (2001) Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers | 0.811 | 4 | 2 | 100% |
| 10 | Chirag Gupta and Aaditya Ramdas (2021) Distribution-free calibration guarantees for histogram binning without sample splitting | 0.737 | 3 | 2 | 100% |
Showing the top 10 of 75 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Calibeating Prediction-Powered Inference | 0.405 | 1 | 1 |