Instrumenting Failure: Scoring Language Models at Every Decision Node in Imperfect-Information Games
Abstract
Agent evaluations report task success and rarely say what the agent did at each decision. The unmeasured step is the belief-action join, where a stated belief becomes an enacted action. We instrument it. Three 12B to 14B instruct LMs play five zero-sum imperfect-information games against four principled opponents. We score every logged decision node against a CFR+ equilibrium reference along eight per-node optimality properties. The failures have shape. A structured share of decisions is blind to the state, and game and model together determine which regime holds. Elicited beliefs sit 0.326 to 0.389 in total variation from the true posterior, and the action best-responds to them on 0.387 to 0.555 of nodes: a belief-action gap read from behaviour. Adding memory moves the constant-action share where it sat at ceiling and reverses direction on Liar's Dice, so the pre-registered aggregate refuses a single reading. Every number carries a certificate: a gate verdict frozen before scoring, and a census that sorts each defect rate to the instrument, a representation, or the models.