She Performs Your Voice: A Unified Speech and Dance Motion Model
Abstract
Expressive performers resonate with audiences because they transform sound into motion: rhythm invites steps, prosody evokes gestures, and energy shapes the flow of the body. However, existing audio-driven character systems often reduce audio to a task-specific condition, separately optimizing for co-speech gestures or music-driven dance. This fragmented view overlooks audio as the shared perceptual driver of performance, making it difficult to generate coherent full-body motion across speaking, dancing, and their intermediate states. We present VOXPERFORMER, a unified audio-to-motion framework that moves toward a generalized performance foundation model. Our key insight is that audio is not merely an external condition, but the hidden soul of performance style, governing how motion emerges, evolves, and transitions. To realize this, VOXPERFORMER introduces a structured motion prior to organize the full-body performance space, audio-to-prior distillation to transform sound into latent motion cues, a memory-augmented prior to resolve audio-motion ambiguity, and a decoupled diffusion generator with implicit transition matching for efficient and coherent synthesis. Experiments on BEAT2 and FineDance show that VOXPERFORMER remains competitive on speech-driven gesture generation while substantially improving hand-aware music-driven dance generation. User studies further show that participants prefer VOXPERFORMER, indicating stronger audio-motion correspondence and more natural perceptual resonance.