ReGen: Agentic Video World Modeling with Synergized Reasoning and Generation
Abstract
Video world models aim to simulate environmental dynamics under interaction. However, most existing approaches entangle transition reasoning with pixel-level generation within a single model, thereby limiting narrative continuity and long-horizon coherence. We reformulate video world modeling as an agentic process that synergizes reasoning and generation, and present ReGen, which couples a streaming vision-language planner with a causal diffusion generator. At each step, the planner infers a structured action-state specification from history, while the generator realizes it as the next video segment. The planner further verifies the outcome and decides to continue, end, or re-generate, forming a closed-loop sense–plan–act–verify pipeline. To support this formulation, we carefully construct a large-scale action-grounded video dataset with over 2M segments that instantiates the planner–generator interface, and introduce action-grounded reward alignment that turns the action specification into an enforced reward and mitigates cross-segment drift. Experiments show that ReGen improves long-horizon narrative coherence over prior work, supporting autonomous long-horizon rollouts without step-by-step user intervention.