Mixture-of-Translators: Scalable Memory Reuse across Heterogeneous LLM Agents
Abstract
Heterogeneous multi-agent reasoning relies on shared contexts, evidence, and dialogue histories, yet agents' key-value (KV) caches remain model-specific. Consequently, every agent must repeatedly prefill or store the same context, so KV-cache memory grows with the number of agents. We propose Mixture-of-Translators (MoT), a cache translation framework that maps context KV caches from a source LLM into a target LLM's cache space. Unlike prior approaches using a single projection path or shared latent space, MoT uses multiple translators with token-level routing. We further introduce a Context Correction Loss that aligns the replayed target trajectory with the native trajectory. Our analysis reveals two competing failure modes: propagation error under early-layer translation and correction-deficit error under late-layer translation. Accordingly, MoT reduces translation mismatch driving propagation error, while Context Correction constrains residual target-side shift from insufficient correction. Across homogeneous and heterogeneous settings involving Qwen2.5, GPT-2, and OPT, MoT preserves downstream QA performance. For heterogeneous translation from a 7B source into a 0.5B target, it achieves 51.0% accuracy and 0.43 F1, recovering 98.1% and 95.6% of native performance, respectively. Enabled by reliable translation, MoT keeps the active KV-cache working set nearly invariant as the number of agents increases while preserving reasoning quality, enabling scalable memory reuse across heterogeneous agents. Our code is available at https://anonymous.4open.science/r/MoT-345F.