Auto-Dreamer: Learning Offline Memory Consolidation for Language Agents
Abstract
Language agents increasingly operate over streams of related tasks, yet existing memory systems struggle to convert accumulated experience into reusable knowledge. Retrieval-augmented and structured memory methods record per-session observations effectively, but often couple acquisition and consolidation into a single online process, leaving the agent without a global view across sessions to discover recurring patterns, abstract shared procedures, or prune redundant entries. Inspired by complementary learning systems theory, we propose Auto-Dreamer, a learned offline consolidator for language-agent memory. Auto-Dreamer decouples fast per-session memory acquisition from slow cross-session consolidation. Given a region of a typed memory bank, the consolidator performs bounded tool-use to search memory, inspect candidate entries, and trace them back to raw source trajectories, synthesizing provenance-grounded replacement memories that supersede the original region. We train Auto-Dreamer via GRPO, using end-to-end agent performance as the reward signal to learn how to consolidate memories acquired through fast online experience. Trained on ScienceWorld trajectories alone, Auto-Dreamer transfers without retraining to held-out ALFWorld and WebArena, improves task success over fixed, RL-trained, and prompted memory baselines in both continual-memory deployment and fixed-bank consolidation, and does so with an active memory bank an order of magnitude smaller than competitive baselines.