WaveGen: Truly End-to-End Waveform Generation via Internal Spectral Trajectory Learning
Sang-Hoon Lee ⋅ Ha-Yeong Choi
Abstract
This work proposes WaveGen, a fully end-to-end diffusion-based text-to-waveform framework based on Spectral Forcing, which jointly learns an internal spectral trajectory and the target waveform trajectory within a single model, enabling direct end-to-end text-to-waveform generation. Unlike recent TTS pipelines, WaveGen does not rely on separately trained or externally decoded intermediate representations, such as neural audio codecs, Mel-spectrogram vocoders, or VAE-based latent models. In addition, WaveGen performs internal text-speech alignment within a single model, eliminating the need for external alignment modules such as duration predictors. To preserve a truly end-to-end formulation, our framework further avoids dependence on self-supervised representations, such as BERT-based text embeddings and Wav2Vec 2.0-based semantic representations, as well as semantic distillation methods for training optimization. To achieve this, we carefully design Spectral Forcing, a model architecture with disentangled diffusion heads, and a training objective based on a multi-scale spectral $v$-loss. We further validate the scalability of end-to-end TTS models by scaling the model size from 0.1B to 0.4B parameters. With a single architecture, open-source data, and one-stage training without external modules, WaveGen demonstrates fully end-to-end waveform generation by achieving promising performance. In particular, WaveGen achieves a WER of 1.77, a SIM of 0.63, and competitive MOS performance on the Seed-en benchmark.
Chat is not available.
Successful Page Load