Supported Facts, Unsupported Inference: Auditing a Drug-Target Prioritization Agent
Abstract
Scientific agents may produce conclusions that agree with external knowledge yet are unsupported by the evidence retrieved in a particular run. In drug-target prioritization, such a conclusion can feed a go/no-go decision. We study this failure in an agent that records tool outputs in an append-only ledger, constructs evidence tables in code, and uses a language model only to route tool calls and generate a verdict with a short rationale. Across nine development checkpoints, the system's guardrails eliminated the observed value drift but did not prevent missing evidential links between retrieved values and conclusions. In an adversarial audit, 27 of the 36 adjudicated allegations were upheld and manually consolidated into 18 defect categories; the validator had flagged none in the original runs. The defects involved external facts or relationships absent from the tool outputs, not fabricated retrieved values. A longitudinal IL6R--coronary-heart-disease example shows that keyword-based checks responded to changes in wording while the underlying unsupported inference remained. We also find that pass rate alone can reward outputs that make few checkable claims. Evaluating scientific agents therefore requires verifying evidence paths for claims and relationships while reporting claim coverage alongside support.