Closing the Reflection Gap: A Free Calibration Bonus for Agentic RL
Abstract
LLMs are increasingly deployed as agents that interact with external environments and observe feedback such as execution results, error messages, and tool outputs. A well-functioning agent should leverage this evidence to assess its own performance. Yet we find that standard RL algorithms systematically undermine this ability: agents drift toward miscalibration, error flags become unreliable, and reflection carries no information beyond raw task accuracy. The root cause is a credit assignment mismatch in outcome-based RL: for instance, relying on outcome alone can penalize honest error detection on failed trajectories. We propose a simple yet effective fix that augments the outcome reward with a free calibration bonus, computed by contrasting the agent's reflection with the actual outcome---requiring no additional reward model, LLM judge, or external annotation. In a text-to-SQL environment across five benchmarks, our method not only improves task accuracy from 75.1\% to 76.5\% but also reduces underconfidence rate from 44.4\% to 7.7\%. The resulting calibrated reflection further enables more effective selective prediction, and supports self-improvement using reflections as pseudo-rewards without outcome supervision.