CacheMAS: Cache Communication for Single-Pass, Jointly-Optimized Multi-Agent Systems
Chao Ouyang ⋅ Yuyang Bai ⋅ Jun Zhang ⋅ Tianlu Gao ⋅ YuxinDai ⋅ xu peidong ⋅ David W Gao
Abstract
Multi-agent LLM systems (MAS) improve reasoning by decomposing across specialized roles, but their text communication is expensive. Each agent decodes thousands of intermediate tokens for the next, and successive prefills grow cumulatively. Recent work shows that latent-space signals can substitute for explicit text between agents, cutting cumulative prefill cost. When the latent reasoning also replaces an agent's text reasoning, it additionally cuts intermediate tokens. Here, we find that the prefill cache alone suffices as the latent-space signal between agents, and the per-agent forward chains can be compressed into a single pass. With \emph{exactly one prefill and one decode per problem}, we sharply cut per-rollout tokens and prefill wall time. We call this architecture \textbf{CacheMAS}, in which prefill agents pass information through cache communication, and a final agent decodes the answer. As a bonus, \textbf{CacheMAS} is one autoregressive rollout, so unlike prior MAS, standard sequence-level RL can jointly optimize all agents at once without per-agent reward design. On math, multi-hop QA, and code, CacheMAS uses $\sim$$3\times$ fewer tokens, runs $\sim$$2\times$ faster, and after joint optimization beats prior MAS by $+1$ to $+8$ points on every benchmark. Cache communication is not just a token-efficient substitute for text, but a practical interface for joint MAS optimization.
Chat is not available.
Successful Page Load