When Does Cache-Aware Context Assembly Pay? A Latency and Quality Frontier for Agent Memory Ordering
Ahmet Erdem Pamuk
Abstract
An agent rebuilds its prompt every turn from retrieved memories, and the ordering of those memories determines how much of the KV cache can be reused. Cache-aware ordering can therefore reduce time to first token, but it may also move relevant evidence to less favorable positions and affect answer quality. We measure this latency–quality trade-off directly using 120 five-turn sessions from a multi-hop question answering workload, controlling turn-to-turn memory overlap at three levels and evaluating thirteen context assembly policies from pure relevance ordering to pure cache reuse. Across 23,400 requests served with vLLM prefix caching, cache-aware ordering reduces median time to first token by 22.4% at high overlap, while providing little benefit when overlap is below approximately 0.2. Answer quality is nearly flat in F1, but exact match decreases monotonically with cache weight at high overlap, corresponding to an estimated trade-off of 8.5 ms of latency improvement per exact-match point, with wide uncertainty. We further calibrate position sensitivity on the same workload and find a within-item slope of only $-0.0011$ F1 per position, indicating that the weak quality cost is partly specific to a regime with little measurable position bias. These results identify a practical churn boundary for cache-aware context assembly and show that its usefulness depends jointly on memory overlap and the model’s sensitivity to evidence position.
Chat is not available.
Successful Page Load