Q/R/M/C: An Intervention-Validated Evaluation Protocol for Memory-Based Agents
Abstract
End-to-end success is too coarse for diagnosing memory-based agents: the same failure score can arise from an unusable query, absent evidence, an incorrect commitment, or failed execution after a correct commitment. We ask when these stage-level diagnoses can be trusted. Q/R/M/C defines four observable constructs— Query formation, Retrieval satisfaction, target Materialization, and Completion— and separates nested query–lock traces, where conditional transition rates multiply exactly to semantic-path completion, from non-nested tool traces, where the same roles are reported as marginals. We evaluate construct, diagnostic, and intervention validity rather than treating the decomposition as a leaderboard metric. On a blinded fault benchmark spanning an OpenAI SDK tool loop, LangGraph, AutoGen, and two local language models, Q/R/M/C attains 0.873 exact-stage accuracy and 0.835 macro-F1 over 63 scored cells. Structural faults are localized in 30/30 cells; the behavioral subset is localized in 20/27 cells; and 5/6 held-out controls are rejected as fault-free. Diagnoses select repairs with mean completion gain 0.293 and regret 0.061 relative to the best available repair. Controlled navigation and multi- agent interventions show when stage profiles track distinct mechanisms, while floor effects, hidden between-stage updates, and missing lock interfaces mark the protocol’s validity boundaries.