Agentic compound selection with dynamic budget-setting
Abstract
Medicinal chemistry campaigns advance in batches of around 20–50 compounds, each batch costing upwards of $100k to make and test. Compound selection is typically driven by unrecorded expert judgement, so a project team cannot learn from its own decisions. Most existing computational prioritisation methods select batch size by subjective human decision. We introduce a digital twin of the compound-selection decision process, rather than a biological system, pairing per-endpoint surrogate models with an LLM agent. We replay two chemical probe medicinal chemistry campaigns (CREBBP, MLLT1/MLLT3) month by month and let an agent set its own per-cycle commitment. Agents reach the same recall–resource frontier as analytic acquisition functions without budget-fraction tuning. Prompt posture sets position along the frontier: justifying which compounds to synthesise rather than exclude changes total synthesis 2.4× while leaving selection efficiency roughly unchanged. Efficiency tracks the quality of the signal provided to the agent. Seeding candidates with expected improvement (EI) recovers 58.0% of the top 50 compounds while synthesising 43.5% as many as the real campaign in CREBBP, and 67.2% while synthesising 32.9% as many in MLLT1/MLLT3. Agent reasoning for every selection is recorded, enabling further integration of human insight and rebuttal. Reasoning traces are grounded in what the agent was shown in 587 of 591 checkable claims.