LLM Is a Good Conditioner: End-to-End Sign Language Video Generation with VQ-Diffusion
Abstract
Sign Language Video Generation (SLVG) aims to generate realistic and motion-accurate sign language videos from spoken text. Most existing studies heavily rely on intermediate representations (\eg poses or 3D meshes), and adopt a two-stage generation paradigm, where the text is first converted into poses or other visual modalities, which are then used as conditions to synthesize sign language video frames via diffusion models (\eg pose-to-image models). However, the end-to-end SLVG task remains largely unexplored. In this paper, we demonstrate that end-to-end SLVG is feasible without relying on intermediate representations(\eg poses), and we argue that a key challenge lies in how to generate effective and temporally consistent conditions. With well-designed condition modules, it becomes feasible to build an end-to-end SLVG model without relying on any intermediate modalities. To this end, we propose LLMCond, a first end-to-end text-to-video SLVG framework that requires no auxiliary visual cues during either training or inference. LLMCond consists of four key components: (1) a video VQVAE that compresses sign language videos into a sequence of discrete latent tokens; (2) \textbf{an LLM-based conditioner} that generates temporally and semantically aligned conditions corresponding to the sign sequences; (3) \textbf{a Gaussian Condition Refiner} designed to enforce temporal smoothness across conditions, enabling the generation of coherent and natural sign language motions; and (4) a discrete diffusion model that synthesizes motion-accurate sign language videos conditioned on the refined condition sequences. Extensive experiments on public SL datasets demonstrate that LLMCond achieves highly competitive performance, producing temporally coherent and accurate sign language motions without relying on auxiliary visual modalities such as pose, depth, or optical flow.