DolphinBench: Evaluating Agent Memory through Action
Abstract
Agents today often take real world actions which depend on long term memory and context recall over time. However, most current memory benchmarks are built for a conversational question-answer format, where the question itself signals that some fact must be retrieved, and often which one. Moreover, benchmarks rarely require anything beyond accuracy from submissions, allowing memory systems to make unreasonable cost/time tradeoffs to achieve higher scores. We present DolphinBench, a benchmark that evaluates memory directly through an agent's task completion. DolphinBench includes three knowledge-work personas with a simulated history of hundreds of sessions and evaluates agents on tasks which depend on information from that history. There are 200 tasks per persona, each of which is verified to be solvable with memory and unsolvable without it. Finally, we require all evaluations to report total cost and latency alongside accuracy, which enables us to evaluate agent memory systems holistically. No existing memory benchmark combines all three. The dataset and evaluation code will be released publicly.