Perception Is Not the Bottleneck for Predicting Human Motion
Abstract
World models increasingly have to reason about the people an agent moves among, and the field treats this as a perception problem: as with scene physics, anticipating human behavior should improve as we observe a person more completely. We show it does not. In a populated social 3D world observed with a rich multimodal telemetry suite (position, gaze, head pose, full-body skeleton, identity, and a social graph), a simple prior over a person’s own recent motion captures what every other channel reveals. Gaze, body pose, whom they face, and their social context are redundant, and a deep model given all of them overfits rather than improves; where a person looks is uninformative even before they move, and the redundancy recurs across worlds. What limits prediction is not sensing but the person’s intent, a latent variable that cannot be observed and must be inferred. Yet intent is not noise: cast as a calibrated belief over goals it is largely recoverable, it sharpens as the person moves, and once it is known the trajectory becomes almost trivial. The lesson for world models of people is a shift from sensing to inference: carry a belief over where others intend to go and keep updating it, rather than perceive them in ever finer detail. Such a belief drives a realistic, world-transferable generative crowd model, and as embodied agents enter human spaces, the ones that share them best will be those that infer intent, not those that see the most.