Identity or Evidence? Auditing Evidence Dependence in Agentic Life-Science Benchmarks
Danai Brilli ⋅ Maria Georgila
Abstract
Agentic life-science benchmarks give an agent evidence about a named entity and score an answer against a recorded outcome. That score tests reasoning over evidence only if the evidence, rather than the name, determines the answer. We examine this distinction by varying identity and evidence independently on a clinical-outcome benchmark and published tasks from Biomni and LAB-Bench. Blinding identity matters little when a prompt contains a decisive sequence or passage. It matters a great deal when the prompt is mostly a name and candidate list (0.742 named versus 0.082 blinded; chance 0.093). Across 96 LAB-Bench LitQA2 questions, the gap increases under one fixed passage-redaction scheme (+0.000 to +0.135; per-instance slope +0.147, $p=0.004$). On the clinical benchmark, a no-evidence recall probe matches prediction accuracy and recovers a database-specific programme count. We argue that benchmarks intended to measure evidence-dependent reasoning should report identity-blinded results, retrieval telemetry, and positive controls.
Chat is not available.
Successful Page Load