When Grounding Fails: Input Verification for Agentic Virtual Screening
Abstract
An agent can execute the intended tools while answering the wrong biological question. We present a retrospective audit of an agentic virtual-screening sys- tem grounded in a medicinal-plant knowledge graph. Its structural annotations in- cluded 10 wrong-protein assignments among 33 registered structures; a historical geometry audit reported insufficient receptor coverage for all ten boxes examined. We introduce deterministic checks for receptor identity, structure-derived search geometry, and pocket preparation within the agent workflow. In the recorded audit, only 9 of 18 target entries remained eligible for docking. Paired controls on two receptors show more negative scores in structure-derived boxes for all 16 tested compounds, with mean shifts of−3.95 and−5.00 kcal/mol. These controls es- tablish sensitivity to input geometry, not binding accuracy or pose recovery. A 210-pair historical campaign and a protocol review expose remaining limitations, including species mismatch and ligand truncation. We argue for evaluating justi- fied abstention alongside task completion, with explicit false-refusal measurement. Appendices report the individual control scores and campaign summaries needed to inspect these claims.