Retrieval Is Not Verification: What a Tool-Using Multi-Agent Drug-Discovery Workflow Actually Looks Up
Abstract
Multi-agent LLM systems are being deployed in professional workflows on the expectation that agents check each other, so that an error is caught inside the system before it reaches a decision. The same exchange carries a false premise from the agent that raises it to the agent that acts on it. Work on that failure usually measures whether agents adopt or recover from the claim; process-level evaluation often only scores whether a tool was called. Neither asks what has to be settled before verification is possible at all: did the evidence that decides the claim ever reach an agent? We make that question answerable. Into a 12-agent target-assessment workflow we inject one externally falsifiable premise as a probe: it names a real database accession and one of its fields, and asserts a value that the record itself contradicts. The tool layer serves biomedical APIs live and freezes each response on first call, so we determine offline, from the recorded tool payloads and with no judge in the loop, whether the evidence needed to verify the claim reached an agent. This workflow retrieves constantly, yet rarely satisfies even that necessary condition. Retrieval reaches at least one exposed agent in 95.8% of runs, the claim's own deciding tool is called in 56.0%, and its deciding evidence arrives in only 23.4%. The bottleneck is not retrieval volume but direction: the two strata diverge by an order of magnitude in where their calls land, and by far less in how many they make. Of answered calls, 57.0% concern the claim's own entity when the claim is about the assessed target, against 4.5% when it is about another gene, and deciding evidence arrives in 52.6% of runs against 11.7%, for the one target these runs assess. The gap persists on a second backbone that retrieves a fifth as much, so neither tracks retrieval volume. A pre-registered matched control separates them directly: a false premise elicits about seven more answered calls per run than a true one, while deciding-tool use and deciding-evidence arrival remain equivalent within the pre-specified margin. Arrival is necessary but not sufficient: an agent can receive a record and still misread it. These rates are therefore upper bounds on verification through the route each claim names, and a workflow that calls a verification tool is not thereby verifying. We will release the instrument and its payload-level scorer, withheld only for anonymous review.