How Agents Know but Fail to Act: Belief Action Gap in Language Model Agents
Abstract
Language model (LM) agents can fail for different reasons, requiring different explanations and interventions. We study how state beliefs, action values, and action preferences interact as LM agents reason and commit to decisions in text-rendered grid worlds. We estimate task-relevant beliefs using report-based and representation-based probes and action values using inverse reinforcement learning. We find that task-relevant state and value information is often represented more reliably than it is reflected in behaviour. Tracking action preferences throughout reasoning further shows that decisions sharpen around a commitment point, belief uncertainty decreases near commitment, and failed decisions commit later on average in a small matched sample, although individual belief changes do not consistently explain shifts in action preference. Finally, forced actions follow patched reasoning trajectories, whereas the tested reasoning-window activation edits provide limited control over the action. Together, our results suggest that some agent failures arise not from missing task-relevant information, but from how available information is translated into and stabilized as an action. These measurements distinguish information availability from action consistency; they do not by themselves identify the cause of each failure.