[Deshraj Yadav] Evaluating the Pareto Frontier of Agent Memory
Abstract
Agents increasingly depend on what they remember across months of work, but memory evaluation has stayed stuck in a question-answer format. A question is itself a retrieval cue: it tells the agent a fact is needed and often which one, so the evaluation begins only after the hardest step has already been done for the system. Real agents acting under an instruction get no such hint. Agent memory should instead be measured through the actions agents take, with two things reported alongside accuracy. Cost and latency, because memory is an optimization problem and a system can buy accuracy by re-reading everything with a frontier model at a price no deployment would pay. And per-test certification in both directions, solvable with the relevant history and unsolvable without it, because audits of widely used benchmarks have turned up wrong answer keys, mislabeled evidence, and judges that accept most deliberately incorrect answers. Early results across built-in memory, dedicated memory plugins, and full context on an open source harness show two plugins reaching nearly identical accuracy while differing by about 70% in cost and close to 3x in latency, exactly the tradeoff that accuracy-only reporting hides.
Speaker