What Do Language Agents Learn by Modeling Environment Responses
Abstract
LLM agents receive rich feedback from their environments, yet multi-turn reinforcement learning typically learns only from final task rewards. Auxiliary environment modeling instead trains agents to predict this feedback, creating an opportunity to learn world models directly from on-policy experience. We study this from two perspectives. Our longitudinal analysis tracks how agent behavior changes throughout training, while our cross-sectional analysis compares how the resulting models predict the same fixed set of environment responses. Late in training, uniform environment modeling leads agents to repeatedly verify tasks they have already solved rather than terminate. These agents continue acting until they exhaust their turn budget because the completed task still receives reward. We test two changes to the environment-modeling objective: annealing its coefficient and giving more weight to moderately surprising tokens. To explain what the resulting models learn about their environment, our cross-sectional analysis uses an autonomous, trajectory-aware pipeline and a shared semantic codebook. We find that most response tokens are already predictable under the pretrained model and that a small high-surprise tail accounts for most of the loss. Although all environment-modeling objectives reduce surprise overall, similar aggregate losses conceal different learning patterns. The objectives differ in how strongly they learn feedback about interaction mistakes, environment inspection, execution problems, and verification outcomes. These findings connect the design of the auxiliary objective to both the world model an agent learns and its behavior during training.