From Weak Data to Strong Policy: Q-Targets Enable Provable In-Context Reinforcement Learning
Abstract
Existing in-context reinforcement learning methods mainly pretrain Transformers with supervised behavior-prediction objectives. While such objectives support task inference from context, they tie the learned policy to the quality of offline actions and can fail when trajectories are weak or suboptimal. We propose Q-Target Pretrained Transformers (QTPT), which preserves context-conditioned inference but replaces behavior cloning with a Bellman-style Q-target objective. QTPT learns to estimate action values from contextual rewards and transitions, rather than simply imitating the behavior policy. We analyze QTPT in stochastic linear bandits and finite-horizon MDPs, deriving suboptimality bounds that separate offline-data effects from Transformer approximation error. Empirically, QTPT is most beneficial under weak offline data across controlled RL benchmarks, and controlled ablations show that context conditioning is essential while Bellman/TD targets provide additional gains. We further include D4RL experiments as higher-dimensional stress tests under matched Transformer baselines.