DuplexGen: Vision-Grounded Full-Duplex Speech Generation via Lightweight Adapters for Joint Audio-Video Diffusion
JIA-HONG HUANG ⋅ Jiajun Fan ⋅ Shuai Wang ⋅ Ge Liu ⋅ Prayag Tiwari ⋅ Stevan Rudinac
Abstract
Full-duplex conversation, where multiple speakers talk simultaneously with natural interruptions and overlapping speech, is a hallmark of human communication. Yet remains beyond the reach of current speech generation systems. Recent work has shown that joint audio-video diffusion models exhibit an \textit{emergent} ability to generate overlapping speech, grounded in the visual co-presence of multiple speakers. However, this emergent capability is limited and fragile: standard LoRA fine-tuning catastrophically collapses multi-speaker generation to single-speaker output for most interaction categories. We propose \textbf{DuplexGen}, a lightweight adapter framework that enhances the full-duplex capability of a frozen joint audio-video diffusion model (LTX-2, 22B parameters) through \textit{Duplex-Aware Cross-Attention} (DACA), which creates explicit speaker-to-audio pathways using spatial video masks. DACA uses LoRA-style adaptation with only 34.6M trainable parameters ($<$0.2\% of the backbone) and is trained on self-generated data without access to the original training set. Across five categories of increasing overlap complexity, DACA improves the speech overlap ratio by 3.3$\times$ on independent simultaneous speech (C4) and 1.4$\times$ on three-speaker scenarios (C5), while consistently preserving multi-speaker generation across all settings. In contrast, standard LoRA collapses to single-speaker output in 4 out of 5 categories. Hyperparameter ablations reveal that DACA is most effective with subtle gating ($\alpha{=}0.01$) and pervasive insertion ($L_\text{skip}{=}2$), suggesting that speaker-aware routing works best as a distributed modulation rather than a strong override.
Chat is not available.
Successful Page Load