SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents
Abstract
Long-horizon action-conditioned video generation aims to synthesize temporally coherent videos that follow complex action instructions over extended horizons. Existing single-shot video generation models typically operate in an open-loop manner, leading to incomplete action execution, hallucinated motions, and temporal drift. To address this, we propose SPIRAL, a closed-loop framework that performs sequential planning and iterative reflection for action-conditioned long-horizon video generation. Specifically, a PlanAgent decomposes a high-level goal into sub-actions that condition video generation, while a CriticAgent evaluates intermediate video segments and provides corrective feedback for iterative refinement. This closed-loop design further supports self-evolving, utilizing planning and verification signals for GRPO-based post-training to enhance the video generator's consistency and action quality over extended horizons. Moreover, we introduce ActVideoGen-Dataset and ActVideoGen-Bench for training and evaluation. Experiments across multiple TI2V backbones with self-evolving show consistent gains on ActVideoGen-Bench and VBench, demonstrating the effectiveness of SPIRAL.