What Must a Cached Agent Decision Observe? An Evidence-Freshness Benchmark for Molecular Workflows
Abstract
We present a reliability benchmark and reference harness for agentic molecular workflows, which reuse remembered decisions while records, models and query results change. We ask what a cached agent decision must observe to remain valid; the evidence concerns cache reliability, not discovery. A trusted interpreter records field reads and exact query scopes, including empty results; for this restricted language, a conditional argument gives equivalence to fresh evaluation. On the eight-commit FreeSolv history, complete scoped checking keeps all 27,006 decisions current with 3,180 refreshes. Post hoc, naive cache baselines fall short: never-refresh makes 801 stale decisions, time-to-live rules 8–801 depending on phase and budget, same-budget random refresh about 628, and dataset-hash invalidation needs 6.1× the refreshes. Checking only the fields read misses absence, but naturally only on one duplicate relationship. In a pre-declared real RDKit release upgrade, cached descriptor-model selections flip in 5 of 18 cells (up to 14% of the top-k) and never-refresh stays stale, but field-only checking suffices there. A pre-declared planted-duplicate control reproduces the gap. Inside an adaptive loop, pre-declared naive caches make invalid admissions that scoped checking avoids. Coherence buys neither utility nor speed. In 240 adaptive campaigns on FreeSolv and Lipophilicity, random acquisition has lower mean final prediction error than the coherent committee controller. Practically, specialized recomputation wins here: it is about 34.1 times faster than scoped checking, and a post-hoc analytic break-even puts the payoff of reuse above about 90.8 µs per fresh decision, against 2.42 µs here. No language model was tested.