Acting without Knowing: Planning, Prediction, and Transfer Dissociate in Interactive Visual Physics
Abstract
Evaluations of large multimodal models (LMMs) often conflate behavioral task success with true physical understanding. However, classical accounts of ``world models'' dictate that physical understanding requires the unified operation of three core capacities: planning, prediction, and positive transfer. Because current benchmarks typically test these three abilities in isolation, they often obscure whether models actually possess a cohesive internal world representation. In this paper, we propose a joint evaluation protocol to test whether LMMs acquire a more complete and consistent understanding of the physical world. Using a suite of virtual physics-based puzzles, models are tasked with placing a tool in a scene to alter its dynamics, iteratively revising their placements based solely on visual feedback. By testing the same models on the exact same tasks, we uncover a stark dissociation between acting and knowing. While LMMs can sometimes solve diverse physical goals under an interactive action-feedback protocol, their success is not accompanied by robust feedback-guided replanning, accurate prediction, or human-like predictive transfer to new actions. Models show weak and inconsistent prediction of the consequences of their own proposed actions, and their predictive transfer to new actions is partial, inconsistent, and less selectively goal-directed than in humans. Ultimately, we demonstrate that the capacities a world model should wholly unify remain only loosely connected. We therefore argue that benchmarks for LMMs should not measure task success alone, but should jointly evaluate planning, prediction, and transfer.