Efficient and Fast Diffusion Language Models with Structured State Space Models
Shweta Verma ⋅ Shreya Pasnoor ⋅ Abhinav Anand ⋅ Mira Mezini
Abstract
Diffusion Language Models (DLMs) offer an alternative to Auto-regressive Models (ARMs) by leveraging bidirectional context and parallel token generation. However, current DLMs are inefficient and require substantial memory because they rely on transformer backbones. This work presents the State Space Diffusion Model (SSDM), which uses a linear time-invariant (LTI) state space architecture to provide a more memory-efficient and faster alternative to Transformer Diffusion Models (TDM). Experiments show that SSDM achieves 24.1\% lower perplexity than TDM, runs 8.3\% faster, and uses 40\% less GPU memory at a 2k context length.
Chat is not available.
Successful Page Load