Error-Guided Bellman Calibration for Semi-Offline Value Estimation
Abstract
We study semi-offline value estimation, where a small on-policy sample supplements a large off-policy dataset. This regime is common in recommender systems and pre-deployment A/B testing, and the challenge is to use the limited on-policy data to correct systematic bias in a given value predictor. Recent work on Bellman calibration (BC) gives a post-hoc, model-agnostic correction along the predicted value. We propose Error-Guided Bellman Calibration (GBC), which trains an error model on the on-policy sample and calibrates the predictor along this learned error axis. We prove a completeness-free refinement guarantee, a theorem showing that the correction adapts to the intrinsic dimension of the prediction bias, and a finite-sample bound on the value-estimation error that isolates our advantage in the refinement term. Across synthetic, CRM, and D4RL benchmarks, GBC improves over existing calibration baselines.