Shared Memory Is Not Primary Evidence: Evaluating Verification and Abstention in Biomedical Agents
Abstract
Multi-agent AI pipelines increasingly reuse one agent’s synthesis as another agent’s input through shared memory rather than repeatedly consulting primary evidence. This creates an evaluation challenge: a low false-claim acceptance rate may reflect either selective verification or simple abstention from unsupported memory. We introduce a controlled biomedical evaluation using 24 source-audited PubMedQA cases, each yielding a clean claim and a single-reversal corrupted counterpart. We evaluate two LLM verifiers under interventions isolating direct memory acceptance, peer-record exposure, displayed provenance cues, and access to the original source. Without source evidence, Qwen showed a strong clean-over-corrupted acceptance gap (23/24 vs. 1/24), consistent with selective acceptance of the paired claims, whereas Gemini returned UNCERTAIN on all 24 clean and all 24 corrupted cases. Providing source evidence increased Gemini’s clean acceptance to 23/23 technically available cases, while Qwen remained near ceiling at 24/24. Peer-record exposure produced no additional ACCEPT events, while the provenance probe showed no positive same-family acceptance signal; both comparisons were floor-limited. These results show why false acceptance alone is insufficient for evaluating shared-memory verification: evaluations should jointly measure false acceptance, clean-claim retention, and behavior when primary evidence is available.