Context Value Informed In-Context Reinforcement Learning
Abstract
In-context reinforcement learning (ICRL) has emerged as an effective paradigm for test time adaptation to unseen tasks without parameter updates. However, existing ICRL methods can exhibit brittle and unstable adaptions, and the mechanisms underlying such adaptions remain poorly understood. We provide a policy-gradient view of ICRL and argue that relying on trajectory-level return feedback can lead to high-variance updates when test time rollout is limited. Motivated by the actor--critic principle, we propose Context Value Informed ICRL (CV-ICRL), which equips the in-context policy with an explicit context value: the expected discounted return conditioned on the current state and accumulated interaction context. CV-ICRL trains a value head and writes its prediction back into the context as a value token, enabling TD-style return targets for lower-variance policy updates. Experiments on the Dark Room, Minigrid, and Procgen testbeds show that CV-ICRL substantially improves the stability of test time adaptation and achieves higher returns across tasks and environments. The source code and data of this paper are available at https://anonymous.4open.science/r/CV-ICRL-D161.