Measurable by Construction: Contradiction Density and the Evaluation of Belief-Forming Memory
Manas Vardhan ⋅ Yug A gupta
Abstract
Existing benchmarks are routinely used to evaluate claims that long-term memory systems revise outdated beliefs. Such a capability can only be tested if a benchmark contains questions whose correct answers require replacing older information with newer, conflicting information. We call the proportion of such questions contradiction density. Measuring it across three benchmarks, we find that an audited classifier returns effectively zero on two widely used benchmarks (0.0% and 0.0–0.3%), whereas model-free derivation gives 81–82% on a third. This difference predicts their ability to detect belief revision: an unchanged supersession mechanism shows null or negligible effects on the two benchmarks with near-zero estimated exposure but a 30-point advantage on the dense benchmark ($p = 1.4 \times 10^{-6}$). An independently developed supersession system shows a larger separation (62% versus 15% accuracy), showing the high score is not unique to a single implementation, while disabling supersession-aware retrieval demotion within our own system reduces accuracy by 16 points ($p = 0.007$). We validate our density classifier against ground truth constructed without model judgments, recovering 91% of contradictions in cases where their existence can be established independently. We further show that integration defects can produce spurious null results, and that self-built evaluation stacks can produce apparent gains that vanish under independently implemented evaluation pipelines. Benchmark composition and integration validation are therefore essential for drawing reliable conclusions about belief revision, and we conclude with a four-step evaluation protocol.
Chat is not available.
Successful Page Load