Hierarchical Context Memory Architecture for Persistent KV-Cache Reuse in Multi-Agent LLM Serving
Abstract
Large language model serving in multi-agent and long-context workflows incur substantial redundant prefill computation because today's serving systems treat each of the inference request as nearly stateless with respect to cross-request or cross-agent context. We present the Hierarchical Context Memory Architecture (HCMA), a three-tier systems design that can coordinates persistent KV-cache and context-state reuse across GPU-resident active memory (Tier 0), a regional in-memory/blob store (Tier 1), and a durable persistent context store (Tier 2), in edge/cloud LLM deployments. A context-aware scheduler routes each request to the lowest-latency tier holding a valid context match and controls promotion and demotion of entries across tiers. Using a trace-driven evaluation on 3,547 replayed conversation requests from the public BurstGPT v2 dataset, HCMA which reduces stateless input-token recomputation from 2,948,306 tokens to 384,088 tokens an 86.97% reduction while adding only 0.250 ms P95 retrieval overhead in the RAM-resident configuration. Discrete-event simulation across ten random seeds at 1,200 concurrent workflows demonstrates 461.51 ms P95 tail latency versus 727.54 ms for a fully stateless baseline (36.5% reduction), with throughput improving from 2.32 to 4.74 req/s (2.04×). A capacity sensitivity sweep shows that approximately 8 MB of Tier-0 storage suffices to hold all the reusable latest-session state for this workload under an 8 bytes/token payload model. We further present component ablation, edge-ratio sensitivity, and context-length sensitivity analyses. Real GPU-cluster validation on A100/H100 hardware is identified as the primary item of future work, and a concrete deployment roadmap is provided.