Unveiling Entropy-Performance Decoupling in Agentic RL for Tool-Integrated Reasoning
Abstract
Tool-Integrated Reasoning (TIR) enables LLMs to solve complex, multi-step problems by offloading computation to external tools. While Agentic Reinforcement Learning (RL) has advanced TIR systems, the underlying policy entropy dynamics governing these models remain poorly understood. In this paper, we conduct an extensive empirical study across model architectures and uncover a pervasive phenomenon of entropy-performance decoupling: in the late stages of RL training, policy entropy continues to diverge despite performance stagnation. We identify that this ineffective entropy growth is primarily driven by an accumulation of invalid tool-use trajectories, which trap the agent in non-informative hallucination loops and degrade final reasoning accuracy. To address this bottleneck, we propose EarlyTIR (Early intervention for Tool-Integrated Reasoning), a strategic truncation mechanism that prunes non-informative exploration during the RL rollout phase. EarlyTIR halts trajectories that exhibit excessive invalid interactions while preserving the model’s ability to learn essential error-recovery behaviors. Empirical evaluations across six public benchmarks and multiple model families (7B to 32B) demonstrate that EarlyTIR consistently elevates the performance ceiling, achieving a macro-average accuracy boost of 13.6\% on the Qwen3-32B model. Furthermore, our approach enhances inference efficiency, reaching correct solutions with 33.7\% fewer tool interactions compared to strong baselines.