MemPilot: Learning Transferable Latent Memory Mechanisms for LLM Reasoning
Abstract
Integrating external memory into Large Language Models (LLMs) typically faces a trade-off between flexibility and depth. Explicit text-based retrieval keeps memory editable, but often acts as a static prompt prefix, increasing context length and interacting with the reasoning process only indirectly. Conversely, parametric memory can influence internal computation more deeply, but updating memory content usually requires additional optimization and may couple domain-specific information with the model's general reasoning behavior. In this work, we introduce MemPilot, a decoupled latent memory framework that dynamically guides the hidden reasoning states of frozen LLMs. MemPilot constructs an external bank of compact latent representations from offline reasoning trajectories, retrieves candidate memory entries, and integrates their latent memory tokens into hidden states via cross-attention and gated residual fusion. By separating domain-specific memory content from the learned memory-use mechanism, MemPilot enables target-domain adaptation through offline memory-bank substitution without updating the base model. Experiments across QA, coding, and mathematical reasoning benchmarks show that MemPilot achieves strong in-domain performance and more robust cross-domain transfer than other memory-augmented baselines. These results suggest that LLMs can benefit from reusable latent memory-use mechanisms while keeping memory content external and replaceable.