Dual-Loop Active Learning for Low-Data De Novo Molecular Generation under CNS and Docking Constraints
Shahzeb Khan ⋅ CHIARA PALLARA ⋅ Isaac Filella-Merce ⋅ Barbara Montini
Abstract
Abstract In drug discovery, designing new molecules for CNS targets is challenging when target-specific data is scarce and desirable properties cannot be reduced to a single reward. We study whether a pretrained chemical language model can be adapted under a constrained regime using nested active learning (AL). Starting with $138$ Salt-Inducible Kinase~3 (SIK3) related molecules, the inner loop filters new chemical entities (NCEs) for drug-likeness, synthetic accessibility (SA), novelty, and BBB-relevant descriptors; an outer loop uses Glide Standard Precision docking to select molecules and write them back into the training corpus. Across six iterations, the model repeatedly improves within fixed constraint levels: property-filter pass rates increase from $54.3\%$ to $60.9\%$ and docking-hit rates from $6.4\%$ to $7.4\%$ in the first iteration pair, with analogous recovery under stricter cutoffs. A fixed-corpus control does not show this recovery, supporting a role for iterative training corpus updating in guiding the search toward the desired chemical space. An alternative Chemformer setup with beam search decoding loses diversity under the same schedule. We therefore present AL as a promising low-data strategy for computationally constrained tasks such as CNS molecular design, while emphasising that the reported evidence is computational and requires experimental validation.
Chat is not available.
Successful Page Load