Task-Aware KV Cache Compression for LLM Agents via Utility-Driven Step Pruning
Yusen Wu ⋅ Yefan Wang ⋅ Jia Yee Tan ⋅ Guangyuan Dong ⋅ Shuang Chen ⋅ Jing Yang ⋅ Rongfeng Guo ⋅ Por L Yee ⋅ Yongtai Liu
Abstract
KV-cache compression for long-horizon LLM agents is often guided by accumulated attention, yet attention frequency can retain repetitive error traces while discarding decisive instructions. We propose TaskKV, a streaming step-level retention policy with three contributions: (1) a signed directional utility label from the first-order change in action likelihood, distinguishing helpful steps from confusing distractors unlike unsigned saliency; (2) a structured residual scorer that treats attention mass and intent alignment as fixed priors and learns only a denoising correction, converging with 500 trajectories; and (3) an entropy-based regime classifier that falls back to contiguous retention when the scorer cannot confidently rank steps. With a 20% KV budget on AgentBench and AlfWorld, TaskKV preserves $\sim$87--89% of full-cache success rate, improving over SnapKV by 5.0--7.9 SR points and over a strong heuristic by 2.9--7.3 points; on dense tasks, gains come from detecting when utility pruning should be disabled. Code is available at https://anonymous.4open.science/r/taskkv-F0D5.
Chat is not available.
Successful Page Load