GeoMemory: Geometry-Indexed Memory for Long-Horizon Interactive Video Generation
Abstract
Autoregressive video diffusion models have demonstrated effectiveness in interactive game generation, with Minecraft gameplay serving as a representative application. To faithfully simulate gameplay, a model must generate natural content when exploring new scenes while preserving spatial consistency when revisiting explored areas. Under limited computation budgets, it must compress and exploit historical cues within a finite context window, which exposes a trade-off: models relying solely on temporal context offer flexible exploration but suffer from poor revisit consistency, whereas adding spatial memory strengthens consistency but may degrade new scene generation quality when the model over-relies on sparse or unreliable spatial context. We present GeoMemory, a learning framework that pairs training protocols with a geometry-indexed spatial memory. Specifically, our Hybrid Training exposes the model to both exploration and revisitation regimes, guiding the model to rely on temporal memory in new scenes while effectively incorporating spatial memory upon revisits. Chained Forward Training creates larger pose variations and encourages reliance on spatial memory for maintaining consistency. For spatial memory, we integrate Point-to-Frame Retrieval with an incremental 3D cache generated by VGGT, enabling constant-time retrieval of relevant historical context regardless of sequence length. Extensive experiments demonstrate that GeoMemory achieves superior performance in both long-term spatial consistency and visual quality in new scenes with real-time interaction.