Cache the Future: Training-Free Self-Revision for Diffusion Transformer Acceleration
Yiyang Li ⋅ Weixuan Huang ⋅ Wei Zhang
Abstract
Although recent training-free acceleration methods have achieved promising speedups for Diffusion Transformers, most rely on historical features, either by directly reusing cached features or by forecasting future features from them. Such a history-to-future design can accumulate errors when the prediction horizon becomes long. To address this issue, we propose FuCa, a Cache-the-Future training-free acceleration framework. Instead of extrapolating future information solely from the past, FuCa caches sparsely computed future features and uses them as online revision signals. FuCa revises the latent-state update and the corresponding velocity, and then recovers skipped timesteps through lightweight interpolation. It requires no extra training or model-specific auxiliary modules, while keeping the number of DiT forward evaluations low. Experiments on FLUX.1-dev, Qwen-Image, Wan2.1-1.3B, and Hunyuan Video show that FuCa achieves up to about 5$\times$ speedup for image generation and more than 4.3$\times$ speedup for video generation, with higher reconstruction quality than existing training-free baselines.
Chat is not available.
Successful Page Load