Behavioural and Mechanistic Evidence of Hidden Goal Inference in AI Agents
Abstract
Can AI agents infer others' hidden goals and act on them strategically? We developed a dynamic, two-player adversarial game where success depends on pursuing one's own goal while inferring and blocking an opponent's hidden goal. We compared the behaviour of five frontier and sub-frontier language models with near-optimal play and a human baseline. GPT-5.6 and Qwen 3.5 outperformed humans partly by inferring goals and blocking strategically. We also fine-tuned a substantially smaller model (Llama-3.1-8B) on near-optimal play. Fine-tuning instilled strategic blocking and produced linearly decodable representations of the opponent's goal posterior and next action. Since the capability to infer and exploit others' goals has direct safety relevance, our paradigm provides a controlled, generalisable way of measuring and interpreting goal inference, revealing strategic behaviour at the level of individual decisions.