Know Where You Stand: Memory-Source Choice in Long-Context Dialogue Agents
Abstract
Memory enables long-context dialogue agents to maintain user state across many turns, but it also introduces a fundamental source-choice problem: grounding the response in the immediate context or retrieved prior memory? A wrong choice incurs two complementary costs: neglecting valid aged-out evidence, or overwriting sufficient context with stale memory. Existing systems and benchmarks largely conflate retrieval with source choice, relying on the model to resolve it implicitly and leaving this critical decision largely unexplored. In this paper, we introduce MSBench, a controlled benchmark that isolates source choice by pairing questions answerable exclusively from context with those requiring prior memory. Under this protocol, strategies that depend on the model's implicit judgment exhibit sharply degraded performance, exposing the limits of leaving source choice ungoverned. To this end, we propose the Memory-Source Reasoner (MSR), which frames source choice as an explicit metacognitive decision: it reasons about contextual sufficiency, answers directly when appropriate, and selectively retrieves missing evidence otherwise. Experimental results show that MSR achieves the best overall accuracy and selection accuracy across four answer backbones. On the primary GPT-5-mini run, it improves over the best non-MSR baselines by 12.6 points in overall accuracy and 6.3 points in selection accuracy. Our benchmark, code, prompts, and evaluation scripts are available at https://anonymous.4open.science/r/msbench-msr-0C72.