SAX: Advancing Video Diffusion Models for Sequential Action Execution
Abstract
Video diffusion models have recently achieved remarkable visual fidelity. However, when handling complex instructions with multi-action sequential prompts, existing models frequently suffer from severe action omission or state freezing. We attribute this bottleneck to two fundamental flaws: i) semantic entanglement, where appearance and motion concepts are heavily intertwined within the text embedding space; and ii) temporal attention leakage, wherein the global nature of cross-attention along the temporal dimension biases the model towards the most visually salient action tokens throughout the entire video. To break this bottleneck, we propose SAX, a novel diffusion transformer framework designed for the synergistic generation of appearance and dynamic actions. First, we introduce a dual-stream attention mechanism that explicitly decouples dynamic action interactions from the static appearance generation process. Besides, to establish a precise temporal-action mapping, we design the Layer-adaptive Action Choreographer. By integrating a shared Global Temporal Table and an Action Ordinal Table, this module dynamically generates a temporal-action alignment map for each layer. Furthermore, to overcome the implicit convergence difficulties inherent in temporal-action alignment optimization, we innovatively employ a Vision-Language Model as a temporal-semantic supervisor. Through distribution distillation, we directly inject the VLM's powerful temporal-action grounding priors into the learning of the alignment maps. Extensive experiments on action generation and prediction demonstrate that SAX substantially outperforms existing models across core metrics.