GraphMemRL: Action-Native Reinforcement Learning for Persistent Graph Memory Construction
Abstract
Long-term memory construction for LLM agents is hindered by two mismatches. The first is a retrieval-objective mismatch: existing systems usually retrieve memories for writing with a single holistic query derived from the current input. However, memory construction depends on several complementary aspects of the input, and a single query is insufficient to capture these editing needs. The second is an optimization-granularity mismatch: long-horizon memory learning is commonly optimized over full trajectories, whereas memory construction proceeds through incremental state transitions. Effective supervision should therefore assess how each local edit changes the memory state. We propose GraphMemRL, a graph-based memory construction framework that addresses both mismatches. GraphMemRL introduces edit-oriented query decomposition to retrieve complementary context for memory writing. It also uses action-native state-transition supervision to directly evaluate whether each graph edit improves the evolving memory state. To make long-horizon learning more tractable, GraphMemRL reformulates memory construction as factorized segment-wise optimization. The agent updates a persistent memory graph across segments under limited context budgets, preserving globally accumulated structure without full-history training. Experiments on long-memory benchmarks show that GraphMemRL improves question answering and produces better connected memory structures with more controlled memory growth. The gains are especially strong on multi-hop, temporal, and knowledge-update tasks.