Towards Efficient Diffusion Language Modeling via Parallel Reservoir Computing
Gabriele Benedetti ⋅ Giacomo Lagomarsini ⋅ Andrea Ceni ⋅ Claudio Gallicchio
Abstract
Diffusion Language Models (dLLMs) have emerged as a promising alternative to autoregressive generation by enabling bidirectional context modeling and flexible iterative refinement. However, dLLMs incur severe computational costs during both training and inference: multi-step denoising over long sequences exacerbates the quadratic complexity $\mathcal{O}(L^2)$ of standard Diffusion Transformers (DiTs). In this preliminary work, we explore an alternative architecture inspired by Reservoir Computing (RC), specifically leveraging Parallel Echo State Networks (ParalESN) as the backbone for diffusion language modeling. By replacing standard attention blocks with parallelized, non-trainable recurrent dynamics, our approach restricts gradient updates to linear projection and conditioning layers, substantially reducing backpropagation overhead while maintaining linear $\mathcal{O}(L)$ inference scaling. We evaluate our method on extremely long sequences ($L=20{,}000$) constructed from the TinyStories dataset. Empirical results demonstrate that ParalESN achieves the fastest training time among all compared recurrent, state-space, and attention baselines (1h 22m vs. 2h 10m for Transformer and 6h 04m for Mamba) with competitive validation loss ($2.47$ vs. $2.38$ for Transformer) and scalable inference latency up to $60{,}000$ tokens. While preliminary, these findings suggest that untrained dynamical backbones offer a promising avenue for scalable and compute-efficient diffusion modeling.
Chat is not available.
Successful Page Load