Sequential Social Inference in Large Language Models
Abstract
Reasoning about others’ mental states requires more than predicting what is true: a model must track how beliefs evolve as evidence unfolds. We study whether process-level supervision can teach this capacity in sequential social inference tasks. To do so, we introduce a controlled ideal observer framework for goal inference, in which a Bayesian observer maintains a posterior over hidden goals under a known rational policy. This provides stepwise posterior targets for evaluating belief tracking as full distributional alignment rather than only as confidence in the true latent state. We pair this with a complementary false-belief setting in which the decisive belief revision event is structurally identifiable. Across both domains, dense reward supervision over unfiltered trajectories can be misleading: ambiguous prefixes dominate the learning signal, encouraging high-entropy predictions that look well calibrated on average but underweight diagnostic updates. Training succeeds when supervision is aligned with diagnostic evidence, either by selecting identifiable trajectories or by applying reward only at the timestep where beliefs diverge. These findings suggest a simple principle for social inverse training: when information is unevenly distributed over time, focus learning on the timesteps most diagnostic of the latent state.