BitSieve: One Bit Budget for KV Selection and Quantization in Block Diffusion LLMs
Gleb Molodtsov ⋅ Ekaterina Alimaskina ⋅ Artur Zagitov ⋅ Evgeny Uskov ⋅ Aleksandr Beznosikov
Abstract
Block diffusion language models decode a block of $B$ tokens over $T$ denoising steps, and their exact prefix KV cache sets both the batch a server can hold and the bytes every block must read. Existing methods cut the reads by selecting a subset of the cache once per block, yet every entry is still stored at 16 bits. We ask where a bit is worth spending. A stored bit costs across all $S$ prefix entries, while a selected entry costs only within the $k \ll S$ entries read once and reused for $T$ steps, and key precision barely changes a top-$k$ decided jointly by $B$ queries. The same bytes therefore buy more accuracy as more entries at fewer bits. We present $\texttt{BitSieve}$, which stores keys and values at 4 bits, ranks the prefix at step 0 directly over the quantized keys, and reads keys by coverage and values by score. On Fast-dLLM-v2-7B and DreamReasoner-8B, $\texttt{BitSieve}$ significantly shrinks the cache, staying close to bf16 selection at the same entry count, reaches the bf16 accuracy plateau at 63\% of its read traffic, and decodes $4.2\times$ faster than dense attention at 28K on one H100. Selection costs 0.28\,ms per block on fused GPU kernels. We release our code at https://anonymous.4open.science/r/BitSieve-Fast-dLLMv2-228D.
Chat is not available.
Successful Page Load