Q-ViK: Question-Guided Visual KV Cache Eviction for Large Vision-Language Models
Abstract
Large vision-language models (LVLMs) rely on increasingly long visual contexts, making the KV cache a major inference-time bottleneck. Existing KV cache compression methods typically use pre-decode signals such as prefill attention, visual redundancy, objectness, or layer-wise allocation, but these criteria do not directly indicate which visual tokens in the KV cache will be reused by future answer tokens. We introduce Q-ViK, a question-conditioned visual utility framework for LVLM KV cache eviction. During offline training, we run full-cache decoding and aggregate answer-to-visual attention over cached visual tokens, yielding a privileged future-utility signal that reflects the visual KV entries actually used during generation. We use this signal to train a lightweight question-conditioned scorer that predicts visual cache utility from prefill representations alone. At inference time, Q-ViK requires neither full-cache lookahead nor generated traces: it preserves textual KV entries and evicts low-utility visual KV entries with a single post-prefill scoring step. Experiments on seen and held-out multimodal benchmarks show that Q-ViK preserves task-relevant visual evidence in the KV cache more effectively than pre-decode saliency, especially under aggressive KV cache compression. Code will be available at https://anonymous.4open.science/r/Q-ViK-1A75/README.md.