DMLLM-Cache: Efficient Caching for Multimodal Generation with Diffusion Language Models
Zhibo Ren ⋅ Guanxi Lu ⋅ Zhican Wang ⋅ Hao Chen ⋅ Wenxi Gao ⋅ Hongxiang Fan
Abstract
Diffusion-based multimodal large language models (DMLLMs) have emerged as a promising paradigm for unified multimodal models (UMMs). However, iterative denoising is computationally expensive and leads to high inference latency. Caching strategies have been explored to accelerate diffusion-based large language models (dLLMs) and diffusion transformers (DiTs), but its potential in DMLLM-based multimodal generation remains underexplored. Multimodal denoising offers greater potential for sparse cross-step reuse than text decoding, but requires caching policies to adapt to different decoding schedules. Motivated by this observation, we introduce DMLLM-Cache, a training-free, plug-and-play caching framework tailored to DMLLM-based multimodal generation. DMLLM-Cache combines schedule-aware adaptive caching with classifier-free guidance (CFG)-guided position selection to identify reusable representations while limiting error accumulation across denoising steps. Experiments on GenEval and DPG-Bench over Lumina-DiMOO and MMaDA demonstrate up to $5.24\times$ speedup with minimal quality degradation and consistent improvements over existing baselines.
Chat is not available.
Successful Page Load