Where Should a World Model Keep the World? 4 KB of Explicit State Beats Both Implicit Memories
Abstract
We train a 3.09M-parameter action-conditioned world model of a Sokoban-like gridworld with an egocentric view, decode it as a per-pixel categorical over tile classes, and play it at interactive rates on a laptop. Because the game is discrete we score the dream exactly, and pair every metric with a control that would expose it if it flattered the model. Aggregate accuracy hides the failure a player sees: at horizon 32 the dream scores 0.750 pixel accuracy, but per-class recall is 0.99 on floor and 0.40 on boxes and goals. The dream empties the level rather than redrawing it, because argmax of a per-pixel marginal returns the modal class and the modal class is floor. Two standard remedies only appear to fix this. Class-balanced cross entropy buys rare-class recall at a far larger cost in precision (at α=1, goal recall 0.396 -> 0.543 but goal precision 0.940 -> 0.031); per-tile sampling barely trades at all, spending precision for a single point of recall. Neither adds information. The missing ingredient is memory: neither implicit mechanism survives a round trip, a K-frame stack forgetting with a hard horizon and a ConvGRU state decaying without one, both falling to 0.63-0.77 tile recall at N=12. Moving static geometry out of the network fixes it. 4 KB of explicit world-coordinate state, with the camera inferred from the model's own frames rather than from privileged state, raises recall at N=12 to 0.992 on every architecture, untrained, at +2.0 ms per frame, while two cheat controls stay flat. A PPO policy trained only in the dream still reaches 88% of real-game return, and the buffer leaves that unchanged, because the task never queries the geometry it repairs, which is a caution about using downstream return as the metric. The question for a small world model is not which implicit memory to use, but whether static geometry belongs inside the network at all.