Transition-Aware Credit Assignment in Agentic Learning for LLM Reasoning
Abstract
Multi-turn agentic reinforcement learning with verifiable rewards enables language models to reason with tools, but standard GRPO-style training often collapses: validation accuracy peaks and then falls while the overall frequency of tool calls remains nearly unchanged. We identify a credit-assignment mechanism behind this failure. With terminal correctness rewards, every tool-related token receives the trajectory’s scalar advantage, making credit transition-blind: the model cannot distinguish useful tool-use steps from unhelpful ones when they share the same final outcome. This transition-blind credit induces two biases. At the token level, failed rollouts with denser tool use can assign negative subset credit to tool tokens. At the trajectory level, token-mean reduction implicitly gives longer failed rollouts larger influence. These biases form a self-sustaining loop that suppresses useful tool interaction without visibly reducing tool-call frequency. To address this, we propose Transition-Aware Credit Assignment (TACA), a novel method that makes tool-token credit depend on the utility of the underlying tool interaction rather than only the terminal trajectory outcome. By restoring transition-sensitive credit and stabilizing sparse tool-use cases, TACA consistently outperforms agentic and non-agentic baselines across different models and multiple reasoning benchmarks.