Why Do LLM Agents Fail to Explore New Environments? A World-Modeling Perspective
Abstract
Single-rollout success can conceal how an LLM agent's behavior changes during learning. We analyze agentic reinforcement learning through Pass@1, which measures success from one sampled trajectory, and Pass@k, which measures whether at least one of k sampled trajectories succeeds and therefore captures broader behavioral coverage. Across five interactive environments, standard RL improves both metrics in familiar settings such as ALFWorld and WebShop. In unfamiliar Sokoban, FrozenLake, and Sudoku environments, however, Pass@k declines while Pass@1 improves only marginally. We call this pattern exploration collapse: policy learning increasingly concentrates on a narrow behavioral mode without acquiring broader environment knowledge. To test this interpretation, we introduce SPA, which collects the base model's self-experience, replaces its state beliefs with true current and next states, and applies supervised finetuning to internalize structured state representations and transition dynamics before policy RL. SPA improves both Pass@1 and Pass@8 across the evaluated small models and tasks. For Qwen2.5-1.5B-Instruct, Sokoban Pass@1 rises from 25.6% to 59.8%. Together, Pass@1, Pass@k, state familiarity, and trajectory evolution provide a practical lens for interpreting agent behavior under environment shift.