EconBase
← All papers

Bellman Calibration for V-Learning in Offline Reinforcement Learning

Lars van der Laan, Nathan Kallus

arXiv 29 Dec 2025 · Statistics — Machine Learning

arXiv:2512.23694 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We introduce Iterated Bellman Calibration, a simple, model-agnostic, post-hoc procedure for calibrating off-policy value predictions in infinite-horizon Markov decision processes. Bellman calibration requires that states with similar predicted long-term returns exhibit one-step returns consistent with the Bellman equation under the target policy. We adapt classical histogram and isotonic calibration to the dynamic, counterfactual setting by repeatedly regressing fitted Bellman targets onto a model's predictions, using a doubly robust pseudo-outcome to handle off-policy data. This yields a one-dimensional fitted value iteration scheme that can be applied to any value estimator. Our analysis provides finite-sample guarantees for both calibration and prediction under weak assumptions, and critically, without requiring Bellman completeness or realizability.

Citation extraction

75
references
128
in-text mentions
75
distinct cited
10
self-citations
6,136
main-text words

appendix boundary found by appendix_command · 38% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Rémi Munos and Csaba Szepesvári (2008) Finite-time bounds for fitted value iteration1.00053100%
2Alexandru Niculescu-Mizil and Rich Caruana (2005) Predicting good probabilities with supervised learning0.92843100%
3Bianca Zadrozny and Charles Elkan (2002) Transforming classifier scores into accurate multiclass probability estimates0.92843100%
4Lars Van Der Laan, Ernesto Ulloa-Pérez, Marco Carone, and Alex Luedtke (2023) Causal isotonic calibration for heterogeneous treatment effects0.9209678%
5Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger (2017) On calibration of modern neural networks0.8434375%
6Jinglin Chen and Nan Jiang (2019) Information-theoretic considerations in batch reinforcement learning0.84333100%
7Chirag Gupta, Aleksandr Podkopaev, and Aaditya Ramdas (2020) Distribution-free binary classification: prediction sets, confidence intervals and calibration0.84333100%
8Tengyang Xie and Nan Jiang (2021) Batch value-function approximation with only realizability0.81142100%
9Bianca Zadrozny and Charles Elkan (2001) Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers0.81142100%
10Chirag Gupta and Aaditya Ramdas (2021) Distribution-free calibration guarantees for histogram binning without sample splitting0.73732100%

Showing the top 10 of 75 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Calibeating Prediction-Powered Inference0.40511