Memory-Efficient Speculative Decoding with Quantized Draft KV Cache
Abstract
Speculative decoding reduces generation cost, but it stores two KV caches: one for the target model and one for the draft model. We reduce this extra memory by quantizing the draft cache. The target cache and verifier remain in BF16, so that compression can not change the target model's final correction rule. We first study how to divide precision between keys and values. A naive per-token quantizer makes keys appear much more sensitive: splitting 12 bits per token, K8V4 beats K4V8 by 10.87%. However, with grouped per-channel keys, the gap becomes -0.25%. A controlled Qwen test confirms that changing only the key quantizer can erase or reverse the preferred bit split. Hence, we propose K4V4 quantization where keys are quantized per-channel while values are quantized per-token. Across six model pairs, our method reduces draft cache memory by 66.13%, and combined memory across both caches by 22.64%. Acceptance only drops relatively by 1.36%. In summary, quantized draft caches can make speculative decoding more memory efficient, but key and value precision must be chosen together with the quantizer layout, as a general predictive principle for efficient AI design.