SceneShifter: Training-free Multi-Scene Temporal Control for Audio-driven Human Animation
Abstract
Although audio-driven human animation has achieved impressive realism, it still lacks effective Multi-Scene Temporal Control: orchestrating background transitions at precise timestamps while preserving a coherent foreground subject. We show that this challenge stems from the diffusion transformer's self-attention mechanism, and addressing it requires two capabilities: (1) Semantic-Aware Temporal Decoupling, which isolates background attention across scenes while retaining cross-scene foreground attention for subject consistency; and (2) Foreground Motion Localization, which accurately tracks the foreground subject across latent frames, especially under large motion. To address these challenges, we introduce SceneShifter, a training-free framework for precise multi-scene control in audio-driven human animation. SceneShifter guides spatio-temporal self-attention by suppressing cross-scene attention among background tokens while preserving cross-scene attention among foreground tokens. To support this guidance under large motion, SceneShifter extracts dynamic foreground masks from the self-attention layers and heads that best capture subject trajectories. We further introduce SceneShifterBench, a benchmark designed to evaluate scene-timing accuracy, foreground preservation, and visual quality under large motion, multiple subjects, and occlusions. Experiments show that SceneShifter achieves frame-level scene timing, strong subject preservation, and high visual quality, outperforming existing baselines. Code and examples can be found at https://anonymous.4open.science/r/SceneShifterneurips26.