MaxIM: Maximally Informative Incremental Summarization via Reinforcement Learning
Abstract
AI systems must often maintain bounded memory of unbounded interaction histories. However, optimizing the memory representation, and its updates, for downstream task utility is difficult for standard reinforcement learning (RL) due to severe credit assignment bottlenecks induced by sparse, delayed rewards. We introduce MaxIM (Maximally Informative Incremental Memory), a framework that formulates incremental summarization as a capacity-constrained, information-theoretic reward-shaping problem. Grounded in value suboptimality bounds, we prove that minimizing sequential information loss requires maximizing the predictive log-likelihood of future, task-relevant targets (predictive sufficiency). We introduce a recursive consistency condition to ensure summaries are grounded by maximizing mutual information with the observed history. We mitigate credit assignment issues with a parameterized critic that computes step-wise variational lower bounds for dense reward shaping. Empirical evaluation on long-horizon benchmarks shows that MaxIM outperforms state-of-the-art baselines on comprehensive autorater metrics and downstream task utility, without expensive human preference labels.