When the Gate Fires: Calibration-Gated Agentic Discovery, and the Structure-Based Methods It Disqualified
Abstract
Agentic discovery pipelines are judged by the shortlists they produce. We report one built so that refusal is a first-class output, and find that on a real programme the refusals were the findings. Three mechanisms carry it: instrument positive controls, so an absence from a corpus that failed its own control is a blocking null; an authority split, in which four of seven modules may only qualify and three may disqualify; and a calibration gate that forbids ranking unmeasured compounds until a method has been shown to track compounds whose potency is known. The limit belongs first: no structural method has passed this gate on either target, and the false-refusal rate is unmeasured. Applied to a 51,479-pair neurological target-disease database, it selected a gain-of-function sodium-channel epilepsy target, where two site transfers of one receptor reach AUC 0.731 and 0.577 against 27 measured actives and 41 measured inactives - measured, not property-matched decoys - with pose controls at 1.57 and 2.82 Angstrom. On a second target chosen because a pass was expected, the redock recovers the crystallographic pose top-ranked and the ranking is still refused at 0.621; on the matched split a 2D-similarity baseline needing no receptor reaches 0.936 where docking reaches 0.602, clearing the floor docking does not. Auditing our own evidence cost four further things and reversed no verdict, including one that would have read QUALIFIED on twelve inactives. We argue that the withdrawal log, the calibration verdict and the audit trail - not the shortlist - are how such systems should be compared.