Fewer Tokens, Fewer Layers: Efficient Vision Token Pruning and On-Policy Distillation to Accelerate VLMs
Shuai Wang ⋅ Shitong Shao ⋅ Qi Xuan ⋅ Zhaowei Zhu ⋅ Jiaheng Wei
Abstract
Vision-language models (VLMs) have achieved remarkable progress. However, processing high-resolution images and long videos generates a massive number of vision tokens, creating a severe inference bottleneck due to the quadratic complexity of attention mechanisms. Previous efforts to mitigate this issue have primarily focused on accelerating the prefill stage by reducing the number of vision tokens, often yielding limited practical speedups while neglecting the increasingly time-consuming decoding stage when the generated response is long. In addition, existing approaches heavily rely on attention maps or similarity maps that are not compatible with efficient attention implementations. To comprehensively address these limitations, we first propose **FlashPruner**, a lightweight module that selects the most helpful vision tokens without relying on attention maps, significantly reducing prefill time. For the decoding stage, we adopt self-speculative decoding and theoretically prove that on-policy distillation guarantees a lower bound for the speculative acceptance rate. Guided by this formulation, we introduce **LoFT**, an on-policy distillation method that effectively trains a highly reliable and efficient draft model including FlashPruner and early layers of the language model. Extensive experiments demonstrate that after training with LoFT, FlashPruner achieves substantial inference speedups while preserving task performance. Notably, on MVBench, FlashPruner improves the time-to-first-token by 1.5$\times$ and decoding speed by 2$\times$, maintaining 96.8\% relative accuracy with only 10\% vision tokens.
Chat is not available.
Successful Page Load