Consistency, Not Plausibility: Dissociating Detection from Correction of Corrupted Tool Records in Biological LLM Agents
Abstract
Agents that answer biological questions through database tools inherit whatever the tools return, and a well-formed record with a wrong value carries no signal that anything is amiss. We ask whether such agents apply biological plausibility checks to tool output, and whether noticing a bad value changes what they answer. An agent answers 40 protein questions with two mock lookup tools backed by real UniProt records; in corrupted conditions one task-relevant field is silently replaced by a value that is either biologically impossible (a 3-residue, 58 kDa protein; a nu- clear human protein attributed to E. coli) or plausible but wrong. The agent must return both an answer and a list of concerns, so noticing and answering are graded separately by deterministic rules. Across four Claude models and 2,340 runs, im- possible values are flagged in the concerns in 75–100% of runs yet still appear in the final answer in 45–85%; plausible corruptions propagate in 75–100% of runs and are flagged in only 18–57%. A severity sweep that scales length or mass by 0.9 down to 0.005, with the paired field either kept consistent or left unchanged, shows what drives the detection that does occur: an internally consistent record propagates at 95–100% at every severity, down to records of 1–5 residues and 100– 600 Da, whereas inconsistent records are partly repaired from the intact field. A rule-based classifier over the 869 flagged concerns confirms the mechanism: flags on inconsistent records cite the contradiction (83%), flags on consistent records cite memory of the specific protein (68%), and only flags that cite a general rule, almost all negative masses, lead to a corrected answer. A one-sentence sanity- check instruction halves propagation of impossible values for the stronger models (odds ratio 0.17 in a task-clustered regression) and reduces plausible propagation by at most 17 points; the whole pattern is unchanged with extended thinking dis- abled. Code, data, traces and grading scripts will be released.