Speak Early: Speculative Speech Generation for Low-Latency SpeechLLMs
Tianyuan Jiang ⋅ Yongheng Deng ⋅ Ju Ren
Abstract
Speech large language models (SpeechLLMs) have emerged as a new interface for human–AI interaction, enabling real-time spoken responses. Time-to-First-Speech (TTFS) is a critical latency metric in such interactive SpeechLLM systems, directly determining perceived responsiveness. While prior work reduces TTFS by accelerating autoregressive decoding, we show that diffusion-based speech synthesis contributes a comparable portion of the latency. This reveals a fundamental limitation of existing approaches that optimize decoding alone, and indicates that optimizing either LLM decoding or speech synthesis in isolation is insufficient. This paper proposes $\texttt{SpecSpeak}$, a new SpeechLLM inference paradigm that overlaps LLM decoding with speech synthesis to reduce the TTFS latency. $\texttt{SpecSpeak}$ introduces a lightweight draft model to generate a prefix quickly. Instead of waiting for fully verified tokens like speculative decoding, $\texttt{SpecSpeak}$ triggers speech synthesis early using draft tokens, while a target model performs parallel verification and correction. To handle inconsistencies, $\texttt{SpecSpeak}$ introduces a verification-aware reuse mechanism that adaptively preserves intermediate diffusion states based on prefix agreement, enabling efficient correction without restarting synthesis. Extensive experiments demonstrate that $\texttt{SpecSpeak}$ reduces the TTFS latency of SpeechLLMs by 18\%--30\%, while preserving response correctness and speech quality. Code will be released upon acceptance.
Chat is not available.
Successful Page Load