CATS: Acceptance-Oriented Critical Token Adaptive Selection for Multimodal Speculative Decoding
Abstract
While Speculative Decoding (SD) has become an essential lossless acceleration technique for Large Language Models, its direct application to Multimodal Large Language Models (MLLMs) is hindered by the intricate visual dependencies present in cross-modal generation. Existing SD methods enforce a uniform global alignment between the draft and target models, which proves ineffective in multimodal settings due to a structured distribution bias: deviations of the draft model are systematically concentrated on a sparse subset of tokens that demand fine-grained visual comprehension. To overcome this limitation, we reformulate the training of multimodal draft models as an acceptance-oriented critical-token alignment problem and introduce CATS, a novel two-stage training framework. CATS employs two complementary selection mechanisms to pinpoint critical tokens:hard tokens that exhibit high rejection probabilities, and vision-critical tokens whose prediction fundamentally depends on visual semantics. Following an initial global alignment phase, CATS performs sparse, targeted refinement exclusively on the union of these identified token sets. Extensive experiments across three representative benchmarks using multiple LLaVA and Qwen2.5-VL target models demonstrate that our approach consistently improves the token acceptance rate over strong baselines such as EAGLE-2 and supervised fine-tuning. Ultimately, CATS substantially accelerates MLLM inference, achieving speedups of up to 3.05×.