Strength Is Not Soundness: Measuring the Exploitability of LLM Policies in Imperfect-Information Games
Benjamin Regazzoni ⋅ Corentin Caillaud ⋅ Rui M Leitão
Abstract
Language models are ranked as strategic agents by strength: average payoff against a chosen field of opponents. Strength is a joint property of the strategy and the field, and certifies nothing against an adversary that models the policy. Exploitability does: the payoff a best response extracts beyond the value of the game, referring to no opponent and computable exactly on small games. It requires the full behavioral strategy, an action distribution at every information set, which play does not provide. Under a frozen black-box protocol we query eight commercial models from four providers at every information set of Kuhn poker ($400$ stateless draws per set, both seats) and best-respond exactly to the resulting table. No hand is played. Exploitability spans $0.111$ to $0.389$ chips per hand, above a Nash-calibrated noise floor of $0.009$. The dominant failure is not choosing the wrong action but emitting no distribution: four of eight models are fully deterministic at temperature $0.7$, and three sit exactly on the optimal deterministic strategy, refusing the randomization that equilibrium requires. Soundness is uncorrelated with field strength ($\tau_b = 0.00$, 95\% CI $[0.00, 0.08]$): equally sound models differ in strength, and equally strong ones differ in soundness.
Chat is not available.
Successful Page Load