Map-Guided Caching: A Global Perspective for Efficient Diffusion Transformer
Jiaqi Ji ⋅ Ran Yang ⋅ Bo Wei ⋅ Hui Li ⋅ Hee Min Choi
Abstract
Feature caching is a promising way to accelerate Diffusion Transformers (DiTs) without modifying the model backbone, but existing methods largely rely on local heuristics or input-agnostic rules when deciding which intermediate features to cache and reuse. We present Map-Guided Caching, a feature caching framework that predicts a Temporal Variation Map (TVMap) for each denoising trajectory and uses the map to derive a global caching policy. The TVMap summarizes block-wise temporal variation across timesteps, while a lightweight Map Predictor estimates this variation from the prompt and initial noise before denoising starts. Given the predicted map, we formulate policy generation as a global path planning problem and combine a dynamic-programming-based orientation planner with a quantity planner that targets a desired speedup ratio. Across image generation, image editing, and video generation, the proposed method achieves up to $3.3\times$ speedup on FLUX.1 [dev] and HunyuanVideo, and $2.7\times$ on FLUX.1 Kontext [dev], while preserving competitive output quality. The advantage is especially pronounced on complex prompts, where local caching policies degrade more sharply. These results suggest that input-adaptive, globally planned caching is an effective alternative to purely local feature reuse policies.
Chat is not available.
Successful Page Load