Imitation Dominates Reinforcement: Direct In-Context RL Is Closer to ICL Than RL
Abstract
In-context reinforcement learning (ICRL) is often described as inference-time RL in which an LLM improves by accumulating trajectory-reward pairs in its context, with the reward acting as a learning signal. This framing poses a central question: is ICRL genuine inference-time RL, or better understood as a form of in-context learning (ICL)? We investigate this in the context of direct ICRL, where the model directly uses trajectory-reward pairs. Through controlled experiments on three benchmarks across six models, we find that the reward is read, but its effect is small; randomizing or removing the reward leaves the improvement curve almost unchanged. Trajectories drive improvement, but not through their semantic content: shuffled or corrupted trajectories work as well as real ones. These patterns closely mirror those known in ICL, suggesting that direct ICRL is better understood as a special case of ICL than as inference-time RL. This reframing has implications for ICRL memory design: ICL factors such as input distribution and demonstrations may matter more than RL elements such as reward shaping and exploration.