Foundation Chemical Language Models for de novo Natural Product Generation and Property Prediction
Ho-Hsuan Wang ⋅ Afnan Sultan ⋅ Andrea Volkamer ⋅ Dietrich Klakow
Abstract
Natural products (NPs), including penicillin, morphine, and quinine, occupy evolution-shaped chemical space and remain prolific drug sources. However, they are largely overlooked by chemical language models (CLMs) widely used for property prediction and molecular design, which were developed mainly for synthetic small molecules. To address this gap, we trained NP-specific CLMs (NPCLMs) on the largest reported NP collection ($\sim$1 million SMILES-encoded molecules), comparing the selective state-space models Mamba and Mamba-2 with a transformer (GPT) baseline across eight tokenization strategies, including character-level, Atom-in-SMILES (AIS), and NP-specific byte-pair encoding (NPBPE). We evaluated de novo generation by validity, uniqueness, and novelty, and three NP-relevant property-prediction tasks—cyclic-peptide membrane permeability, taste, and anti-cancer activity—using Matthews correlation coefficient (MCC) and AUC-ROC. Under random splitting, the state-space models generated slightly more valid and unique structures with fewer long-range structural errors and modestly outperformed GPT in property prediction, whereas GPT produced more novel scaffolds. Under stricter scaffold splitting, which groups molecules by core structure, all models performed comparably, while chemically informed tokenization consistently improved performance. Finally, NPCLMs trained on only $\sim$1 million NPs matched the general-domain ChemBERTa-2 and MoLFormer models trained on over 100 times more data, highlighting the value of curated, domain-specific data over sheer scale for navigating NP chemical space.
Chat is not available.
Successful Page Load