MFlowAudio: Efficient Text-to-Audio Synthesis via Mamba-based Stateful Flow Matching
Hao Dai ⋅ Panyu Chen ⋅ Jagmohan Chauhan
Abstract
Recent advancements in audio generation have been largely driven by Transformer-based diffusion models. However, these models suffer from quadratic complexity of self-attention, severely bottlenecking the long-form audio synthesis. To overcome this limitation, we propose MFlowAudio, a novel latent audio generation framework that synergizes the continuous-time dynamics of Flow Matching with a custom-designed TFMamba backbone. TFMamba uses an innovative dual-scan mechanism: a TimeMamba module to capture long-range causal dependencies with linear complexity, and a FrequencyMamba module to model spectral correlations such as harmonic structures. Exploiting this structural foundation, we formulate the Stateful Flow Matching (SFM) paradigm. This framework inherently enables chunk-wise training and streaming generation, maintaining an $\mathcal{O}(1)$ caching complexity without incurring extra computational overhead. To enable fine-grained controllable synthesis, we devise a novel guidance mechanism that neutralizes vector field collisions precipitated by on-the-fly semantic transitions of prompts. Comprehensive empirical evaluations confirm that MFlowAudio yields generative fidelity on par with state-of-the-art baselines while establishing significant computational efficiency, achieving an $81.8$% acceleration in generation and facilitating perceptually seamless streaming synthesis with a constant, ultra-low latency of $1.8$ seconds. Demo:https://huggingface.co/spaces/mflowaudio/MFlowAudio
Chat is not available.
Successful Page Load