EAT: Eviction-Aware Training for Long-Context LLM Inference on Edge Devices
Abstract
Deploying large language models on edge devices requires mitigating two major bottlenecks: the computational cost of the model and the memory overhead of the KV cache. While combining model quantization and KV-cache eviction seems like a natural solution, we reveal that this naive pipeline suffers from a critical failure mode. The issue stems from a hidden conflict: modern eviction methods rely on attention scores to identify important tokens, yet quantization inherently adds noise to these exact scores. Consequently, quantization does not merely degrade the numerical precision of the KV cache—it actively alters the eviction decisions, often causing the model to discard crucial context. To overcome this non-orthogonal degradation, we propose EAT, a training framework that integrates discrete eviction behavior into quantization-aware training. EAT computes eviction from quantized attention, applies the resulting mask in the forward pass, and uses a masked straight-through estimator in the backward pass. The method requires no architectural changes and adds no inference-time parameters. Systematic evaluations show that EAT mitigates the failure mode of naive quantization-plus-eviction pipelines. Under W4A16 + KV-INT8 with 50% KV eviction, eviction-unaware composition reduces the five-task average from 38.37 to 29.04, whereas EAT recovers it to 39.44 under the same cache budget. Real-world deployment on the MediaTek Dimensity 9500 demonstrates that, under a 50% KV-eviction budget, EAT improves decoding throughput by 64.1% and reduces total memory traffic by 20.4% compared to quantized full-KV inference. Mechanistic analyses suggest that the improvement is associated with reduced attention sink behavior rather than with sharper attention.