How and What Does a World Model Remember?
Abstract
When a player walks back through a scene, a real-time interactive world model redraws paths, objects and rooms it can no longer see. Matrix-Game 3.0 credits this consistency to an explicit memory stack: a bank of latents retrieved by camera-frustum overlap, one latent pinned from the first view, and Plücker-style camera conditioning. We dissect the released model at inference, switching each mechanism off with frozen weights and scoring the result with a mirror test: walk out, turn around, walk home, and check whether the model draws the same hallway on the way back. We score only frame pairs beyond the short-term window, under paired seeds and cluster-bootstrap confidence intervals. The memory bank helps, and its value grows with loop distance, but most of its value needs no retrieval at all: random slots recover most of the effect, the pinned first view recovers over half by itself, and content-similarity retrieval is worse than random. Camera conditioning is asymmetric: the model degrades gracefully when geometry is absent and far worse when its memory registration is corrupted. At these horizons, what the model chiefly remembers is its last sixteen frames and a postcard of home. At 320-frame loops, geometric selection finally beats random, and one implication survives testing: eight retrieval slots instead of five, a two-line change, improves the released configuration by −0.022 LPIPS.