Diversity by Construction: Grounded Question Generation for Agentic Commerce Memory
Abstract
Most benchmarks for agentic memory score a final answer produced over synthetic conversation histories. This design conflates storage, retrieval, and reasoning failures, lets long context masquerade as memory, and rarely tests abstention. We present a compositional evaluation framework that generates evidence-grounded memory questions directly from \emph{real} customer order histories. Sampled semantic axes (memory function, memory demand, predicate, target, and ordering condition) define what each question tests before any language is written. Account data only fills the open slots. Diversity is treated as a measured property with three components: \emph{variety} (linguistic and surface-pattern spread, enforced during diversification), \emph{coverage} (which categories of each axis are sampled, drafted, and exported, tracked cell by cell), and \emph{complexity} (how many axes constrain a question and how demanding its memory operation is, from single-fact retrieval to multi-step inference). A pipeline drafts, validates, diversifies, answers, audits, and dual-judges each question, and a classification flywheel checks that each exported question still tests what its specification requested, flagging mismatches as concrete taxonomy fixes. On a 100-account run, 4{,}000 sampled specifications yielded 2{,}218 exported questions with 99.6\% of constrained axes matching their request, complete marginal coverage of the enabled taxonomy, and 13.8\% unanswerable questions, whose correct response is to abstain because no supporting evidence exists. We use seven exported-content configurations derived from three memory systems and raw order history as a diagnostic case study. The configuration with the highest answerable-question accuracy (67.8\%) also has the lowest correct-abstention rate (52.8\%), exposing a retention--calibration trade-off. These snapshot-level experiments measure information available to a fixed reader, not production retrieval or end-to-end memory-system quality.