A Unified Audio Language Model with Text-Aligned Factorized Audio Tokenization
Abstract
We present UniALM, a unified audio foundation model built on FactorCodec, a purely discrete tokenizer that factorizes audio into two role-separated representations. Analysis tokens are optimized to retain language-aligned, text-expressible information that can be decoded by a text LLM head into grounded natural-language analyses, while reconstruction tokens preserve the information required for waveform synthesis and are used exclusively for audio decoding. This factorization yields understanding performance competitive with continuous Whisper features on audio understanding tasks, while reducing the perplexity of reconstruction-token modeling for generation. UniALM further adopts functional layer specialization, partitioning the backbone into audio-understanding, cross-modal, and audio-generation experts. We train the model with a four-stage recipe on 100B text tokens and 60B audio tokens, together with a composed audio sequence construction strategy for unified multi-task pre-training. At 3B parameters, UniALM is competitive with strong 7B unified baselines on in-domain tasks and exhibits non-trivial \emph{compositional generalization} to unseen tasks under few-shot and zero-shot evaluation. Unlike prior unified systems that rely on hybrid continuous-discrete inputs or do not support generation beyond speech, UniALM provides a purely discrete interface for both understanding and generation across speech, sound, and music.