Optimistic Q-value Adaptation for Offline-to-Online Reinforcement Learning
Abstract
Offline-to-online reinforcement learning (O2O RL) aims to improve offline RL agents with limited online interaction. Prior work has shown that, in conservative value-regularized methods such as CQL, pessimistic offline critics can underestimate values and hinder online fine-tuning. We show that this issue is more general: critic underestimation also appears in explicit policy-constraint methods such as TD3+BC, and can persist or become more severe through bootstrapped Bellman updates during fine-tuning. To address this problem, we propose Optimistic Q-value Adaptation (OQA), a framework that corrects underestimation at the Bellman-backup level. OQA constructs an additional offline-pretrained actor--critic pair and uses it together with the original pair to form a controlled optimistic backup that mitigates inherited pessimism during fine-tuning. Unlike trajectory-return calibration methods, OQA does not rely on return lower bounds and is not tied to CQL-style regularization. Across D4RL MuJoCo locomotion, Maze2D, AntMaze, and Adroit benchmarks, OQA instantiations based on TD3+BC and CQL consistently outperform the corresponding backbones and strong O2O RL baselines.