From Model Scores to Defensible Agent Decisions: Decision Contracts for Molecular Generation and Docking
Hai Luong
Abstract
Computational drug-discovery pipelines chain generation, ligand preparation, docking, and ranking, but components are usually benchmarked in isolation and preprocessing is often treated as deterministic. We introduce DecisionBench-DD v1.0, a frozen, auditable decision-contract methodology that scores tools by end-to-end candidate yield under task-specific gates, and extend that contract to the LLM agents that increasingly interpret such results. In an ambroxol-conditioned GCase/GBA1 case study, 66/300 GenMol and 3/300 MolMIM outputs passed chemical qualification (configuration–contract fit, not generator superiority). A contract-aware purchasable-catalogue retrieval policy instead returned 258/300 qualifiers after searching 9.32 million unique molecules, only 804 of which fell in the required seed-similarity band— a configured acquisition-policy comparison, not a symmetric algorithm or compute comparison (Appendix J). In redocking, Top-1 was unresolved (3/14 vs. 2/14), but GNINA recovered more near-native poses within five predictions (8/14 vs. 2/14) and produced more geometry-valid poses. Fresh embedding changed 18/69 candidate-screen labels and 3/14 redocking verdicts under a fixed seed, and a property-matched decoy comparison gave an inconclusive risk ratio of 1.46 (95% CI 0.96–2.12); the 16 frozen candidates are an auditable prioritization set, not validated hits. We then tested whether the same framework's evidence-semantics rules change how an LLM agent interprets this evidence: across three prompting conditions (free-form, prose guidance, explicit numbered contract) replicated 50 times (weaker model) and 5 times (stronger model), nine of ten items were answered correctly in every run. The exception—overturning a null primary endpoint with an attractive secondary one—showed the weaker model's violation rate fall from 64% (free-form) to 34% (prose; $p=0.005$) to 28% (contract; not significant vs. prose, $p=0.67$), while the stronger model was correct in all 15 runs but only the contract condition left a rule-cited audit trail: stating the rules at all helps most; formalizing them into a contract adds auditability without a further, resolved accuracy gain. Decision contracts should be frozen to compare tools honestly and to constrain the agents that increasingly sit between pipeline outputs and advancement decisions, while conformer preparation should be treated as an experimental factor and structural proxies calibrated before compounds advance.
Chat is not available.
Successful Page Load