OopsWorld! Operation-Grounded Seamless World Generation
Abstract
We study seamless world generation, where a model must update a persistent visual world according to online user operations rather than full next-scene descriptions. To make this task well-defined, we introduce an operation-centric taxonomy that represents each scene with a global-local schema and each boundary with a typed operator program specifying change and preservation. This taxonomy enables SeamlessWorld, a synthetic data and benchmark pipeline with 8 operation families, 23 atomic operators, dual prompt surfaces, and schema-derived evaluation probes. We then develop OopsWorld, a causal video generator that learns from SeamlessWorld using operation-conditioned decoupled distillation: the student is driven by operation prompts and visual history, while scene-conditioned teachers provide target-state supervision. We further apply dynamics curriculum learning to progressively compose simple edits into high-dynamic world transitions. OopsWorld substantially improves dynamic responsiveness and long-video quality, achieving 91.67 VBench Dynamics and 80.42 Long-Quality with 100K samples. A systematic explicit--implicit study further reveals a persistent operation-prompting gap in current systems, especially for preserving visual anchors, calling for operation-driven world models beyond next-scene prompt switching.