Assessing Scientific Claims in Agentic Workflows: MAER and GAVEL-X for Biomedical Research Evidence
Abstract
A biomedical-agent answer that passes a benchmark does not by itself establish what the evidence justifies. This paper develops Multi-Agent Ecosystem Risk (MAER) and GAVEL-X (Governance and Audit for Verified, explainable LLM systems) to assess biomedical research evidence against an explicit claim and its intended use. MAER specifies evidence requirements for claims about interactions; GAVEL-X links judgments to executions, scientific requirements, sources, and revisions. Revision rules distinguish recovering evidence, changing requirements, rerunning a computation, revising a claim, and changing the evaluator. An audit of two BiomniBench-AI4S score representations finds that two judges disagree on acceptance for 52 and 45 of the same 350 backend–task pairs, respectively; neither count estimates biological error. In 48 scripted numerical workflows, each of four combinations of assessment procedure and record format reaches justified conclusions on 168 of 240 claim–evidence-view pairs and leaves 72 unresolved. When decisive records are available, obligation-guided recovery and a competent dependency checklist each resolve all 24 initially unresolved cases; static review resolves none, and uniform random retrieval resolves 22.22% in exact expectation under the same two-request limit. These results test software conformance and recovery of existing records, not human or language-model assessment performance. The method distinguishes evidence about a measured scientific process from records of its analysis. The contribution is a versioned assessment method with executable component tests and a protocol for future assessment studies on biomedical research records. A separate, proposed interface would connect assessment to longitudinal investigations using numerical models. No human or language-model assessment study or integrated investigation has been run. Assessment benefit and discovery acceleration therefore remain untested.