Verify, Rebind, Repair: Enforcing Evidence Consistency in Small Biomedical Agents
Abstract
Agents that propose drug-repurposing mechanisms can cite valid evidence while making an unsupported claim. We study evidence contracts that let deterministic tools check whether an answer is consistent with supplied evidence cards. On 1,280 DrugMechDB graph tasks per model over 256 held-out drugs, citation rebinding and abstention filtering raise success by 26.8–86.6 points for three small models without another LLM call, mainly by recovering omitted citations and enforcing abstention. In a post-hoc replay, escalating only answers without a verified path to a 27B model lifts all three pipelines above that model with the same tools (Qwen3-8B: 99.5% vs. 97.2%, +2.3 points, 95% interval [1.4, 3.1], with 23% of tasks escalated), and by construction it never scores below it. Violation feedback beats self-reflection for two of three models. On DrugProt abstracts, the prespecified comparison of specialist repair with matched reflection is inconclusive (+1.6 [-0.8, 3.9] and +2.0 [-1.6, 5.5] points). Secondary and exploratory ablations favour the specialist's label for Phi-4-mini (+7.8 points over direct generation, and +8.6 for the label alone over withheld feedback). A content-blind excerpt passes the literal quotation check, so text evidence needs entailment checks.