Grounding Agent Reasoning with Structured Process Supervision for Multi-turn Reinforcement Learning
Abstract
Long-horizon LLM agents trained with outcome-based RL exhibit a systematic failure mode: reasoning diverges from environmental observations at some step, and the error compounds across subsequent turns. We find that 74\% of failed trajectories contain reasoning–observation inconsistency, and once it appears, 62\% of subsequent turns remain inconsistent. Process rewards reduce but do not eliminate this: under free-form reasoning, agents learn to avoid penalties without genuinely tracking task progress or anticipating action outcomes. We propose ARC (Anchored Reasoning with Commitment), which requires the agent to make two explicit commitments at each step: a global progress statement that reconciles the agent's understanding of task state with accumulated observations, and a local outcome prediction that must be borne out by the next observation. Together, these commitments anchor the agent's reasoning to the actual environment state across turns, and provide well-defined targets for process supervision. To prevent interference between sparse outcome and dense process signals, we decouple their advantages via independent group normalization. On ALFWorld, WebShop, and WebArena, ARC improves over the strongest baseline by 6.4, 3.5, and 9.6 points respectively.