Accelerated Image Editing via Consistency-Aware Source Token Pruning
Jeongsol Kim ⋅ Jong Chul Ye ⋅ Yufei Wang ⋅ Jian Wang
Abstract
Text-based image editing has recently been reinterpreted in large multimodal transformers as conditional generation, where source image tokens are concatenated with text and noise tokens. While effective, this design incurs substantial computational overhead in attention layers. To mitigate this inefficiency, we propose {\em TokenDrop}, an efficient editing framework that shift the computational burden from dense source-token conditioning to a lightweight regularized sampling update. By transferring the influence of source tokens into a closed-form sampling update, the method preserves source consistency with negligible cost. To support this regularized dynamics, we introduce a deviation-guided adaptive masking strategy that selectively drops redundant tokens while maintaining editing fidelity. Across FluxKontext and Qwen-Image-Edit, our training-free method achieves an average 22.4\% improvement in inference speed on PIEBench, while better preserving non-edited regions. The method delivers up to 1.8$\times$ speedup at 1024$^2$ resolution and 2$\times$ speedup at 2048$^2$ resolution. The code will also be released publicly.
Chat is not available.
Successful Page Load