Intent-Aware Caching for Efficient LLM Serving
Abstract
Modern LLM serving systems use prefix caching to accelerate multi-turn conversational workloads. They cache the key-value states of previous conversation prefixes and reuse them when subsequent requests share the same prefix, to avoid redundant prefill computation and reduce serving latency. However, achieving a high cache hit rate remains challenging because cache reuse depends on whether a conversation continues and how soon the next turn arrives. Our analysis of real workloads shows that user intent is a strong signal for this behavior, as different intents exhibit different continuation probabilities and inter-turn delays. Based on this insight, we propose Intent-Aware Caching (IAC), an online eviction policy that classifies requests by intent, learns lightweight reuse statistics from recent serving logs, and uses them to guide an oracle-inspired eviction decision. Extensive evaluation on real workloads (e.g., ShareChat) shows that IAC reaches a 34.6\% cache hit rate, substantially outperforming the ML-based LPC (18.8\%) and LRU (12.2\%), while reducing average serving latency by up to 54.6\%.