D2ML: Models for Direct-to-Biology Small Molecule Screening
AkshatKumar Nigam ⋅ Neri Amara ⋅ Hillary J Dequina ⋅ Dyana N Kenanova ⋅ Seth Malmersjö ⋅ Fiorella Ruggiu ⋅ Raymond V Fucini ⋅ Jack Sadowsky ⋅ Stig K Hansen ⋅ Imran S Haque
Abstract
Molecular property prediction is a central problem in machine learning for drug discovery. Property prediction models, such as those for protein–ligand binding prediction, are often highly data-limited, with predictive performance degrading substantially as molecules move outside the training distribution (Wallach and Heifets, 2018; Fooladi et al., 2025). This challenge is compounded by the structure of chemical property landscapes: most molecules may be inactive for a given problem, while compounds that are nearby in chemical space can nevertheless differ substantially in activity, a phenomenon known as an “activity cliff” (Maggiora, 2006; Stumpfe and Bajorath, 2012). Small-molecule discovery requires searching an enormous chemical space under economic constraints: custom chemical synthesis costs on the order of US\$1000 per compound, and each design cycle takes on the order of 10 weeks. As a consequence, conventional search operates in two phases. Hit identification uses methods such as high-throughput screening or DNA-encoded libraries, which amortize pre-established libraries across screens of multiple targets, to broadly sample chemical space and identify initial hits. In hit-to-lead and lead optimization, medicinal chemists cyclically refine these hits by synthesizing a small number of close analogs to explore local structure–activity relationships while limiting experimental cost. As a result, activity datasets often suffer from a “streetlight effect,” in which active compounds occur in closely related series—not because they are the only possible actives, but because they are particularly accessible through local search (Blevins and Quigley, 2025). Models must therefore learn from data that are scarce, chemically biased, and subject to temporal distribution shift as discovery programs progress. Useful models must operate within a design–make–test–learn cycle in which experimental resources are limited and compounds are evaluated against multiple objectives rather than a single endpoint. An alternative strategy combines nanogram-scale high-throughput chemical synthesis with biological testing of crude reaction mixtures—that is, taking reactions “direct to biology” without purification—allowing $>10^4$ novel compounds to be made and tested in days. Carmot Therapeutics was the first to industrialize this high-throughput chemical synthesis and direct-to-biology strategy, abbreviated D2B, as a central approach to chemical exploration, successfully discovering multiple drugs using an empirically driven rather than model-guided process (Hansen et al., 2018; Shin et al., 2019; Rodriguez et al., 2025). Extending this earlier empirical strategy toward model-guided D2B introduces new challenges. Even in conventional chemical search, successful modeling must handle variable-fidelity data from high-throughput screening and focused low-throughput assays; to benefit from D2B’s scale, models must additionally account for the complexity introduced by testing unpurified compounds with uncertain true concentrations. In this work, we investigate how molecular property prediction models can make use of the structured supervision produced by a D2B assay sequence. Rather than treating measurements from different assay stages as interchangeable labels, we match the prediction task to the information provided by each experimental readout and investigate whether information can be shared across stages to improve prediction where high-fidelity measurements are scarce. We evaluate this approach prospectively in a drug-discovery setting and demonstrate simultaneous improvements in potency and selectivity, showing its application to the exploration of a large, synthetically accessible chemical space.
Chat is not available.
Successful Page Load