Information-Directed Offline-to-Online Reinforcement Learning
Keru Chen
Abstract
Offline-to-online reinforcement learning first warm-starts a policy from a fixed offline dataset and then improves it with limited online interaction. Offline data reduces uncertainty, but it does not remove the need for exploration; it changes what remains to be explored. We formalise this residual uncertainty by the conditional mutual information \$I(\chi;\tau_{1:T}\mid\mathcal\{D\}_N)\$ between a learning target \$\chi\$ and the online trajectories after conditioning on the offline dataset. This view leads naturally to information-directed sampling (IDS), a family parameterised by \$\eta\ge 0\$ that selects actions by trading off instantaneous regret against information gain. We prove a generic offline-to-online Bayesian regret bound for IDS through a ratio certificate: any information-ratio bound satisfied by a reference Thompson-sampling policy over the same randomised policy class is inherited by IDS. In a known-dynamics Bayesian linear-reward model, the conditional mutual information has a log-determinant form, and vanilla IDS (\$\eta=0\$) satisfies \$\widetilde O\(Hd\min\{\sqrt T,\,T\sqrt\{C^\dagger\_{\beta,\mathrm\{IDS\}_0}(N,T)/N\}\}\)\$, where the coverage coefficient is tied to the visitation distribution induced by vanilla IDS itself. We also identify a warm-start regime with a dominated but informative probe in which vanilla IDS selects the probe while Thompson sampling never does, giving a constant-factor Bayesian regret separation. Controlled bandit experiments and D4RL offline-to-online experiments support this mechanism: IDS is most beneficial when offline data is informative but leaves biased or low-probability residual uncertainty that can be resolved by targeted online actions.
Chat is not available.
Successful Page Load