When Behavioral Signal Is Not Enough: Separating Observability from Learned Theory-of-Mind Inference
Abstract
Evaluations of Theory of Mind (ToM) in large language model (LLM) agents usually report a single inference score, which conflates two different failures: the opponent’s behavior may carry no signal about its hidden state, or the signal may be present and the model may fail to extract it. We separate these in a controlled negotiation testbed with nine executable opponent types, where an exact, mechanism-matched observer measures how much information the behavior actually carries, and a learned LLM observer is scored against it on the same traces. The two dimensions of the opponent’s latent state come apart sharply: the exact observer finds substantial signal for both dimensions on their respective prespecified scopes, while the learned observer captures 64% of the available orientation signal but only 14% of the available depth signal. We argue that behavioral observability and learned recoverability should be measured separately.