Unbounded Streaming Text-To-Speech with Prefixed Sliding Window Attention
Abstract
Existing text-to-speech (TTS) systems achieve high quality, yet streaming and long-form generation remain challenging. We propose simple adaptations to a standard encoder–decoder TTS architecture that enables both capabilities without limits on total duration. We first observe that attention patterns in these models are well organized, following a structure that motivates a specific windowing strategy: a prefixed sliding window for autoregressive decoding and a sliding cross-attention for text conditioning. Together, we show that these adaptations address length generalization and error accumulation while being fully streamable in text and audio. Trained exclusively on segments shorter than 30 seconds, our system produces seamless, consistent speech at practically unbounded lengths with linear decoding complexity, outperforming baselines that rely on hour-long training context. We validate the approach through continuous synthesis over hours long synthesis and demonstrate competitive results against state-of-the-art open-source baselines on both short- and long-form benchmarks.