Do multimodal models imagine electric sheep?
Abstract
Yes. We find that large multimodal models develop mental imagery when solving spatial puzzles, and they do imagine sheep when solving sheep puzzles. We fine- tune a Qwen3.5 VLM to solve twelve diverse visual reasoning tasks—including tangram, jigsaw, sokoban, 3D mental rotation, and rush hour—that require under- standing geometry, spatial constraints, and the consequences of actions. By simply predicting the open-loop sequence of actions to solve a puzzle from an initial state, we show that the model learns successfully, and that its hidden representations after each action encode meaningful visual information about the intermediate puzzle state. This finding suggests that an imperfect visual world model begins to form as a byproduct of learning to select correct actions in an open-loop fashion, in the absence of any explicit visual supervison. Building on this observation, we propose two ways to sharpen the mental images learned by the model through explicit visual supervision or explicit visual tokens. Across puzzles, we find that integrating as few as sixteen visual tokens into the chain of thought per board state improves the average solve rate from 83% to 89%, with gains exceeding 20 percentage points on reasoning-heavy games such as jigsaw and 3D mental rotation. Our results show that visual reasoning is a bottleneck, and that mental imagery may be a key mechanism for robust spatial cognition.