SIGMol: Enabling Molecular Design with Flow Map Chemical Language Models
Steven Dunne ⋅ Sebastian Ibarraran ⋅ Frank Hu
Abstract
Property-directed molecular design is a pivotal step in drug discovery and lead optimization campaigns, but it is difficult in practice due to the combinatorial growth of chemical space. Chemical language models (CLMs) trained on large corpora of molecules have proven to be a powerful generative modeling paradigm for navigating this vast design space and proposing viable candidates. However, this class of models suffers from a key limitation in that most chemical language models do not leverage reward-based inference-time scaling and steering methods, opting instead for additional post-training via supervised fine-tuning or reinforcement learning. Here we introduce $\textbf{S}$teered $\textbf{I}$nference $\textbf{G}$eneration for $\textbf{Mol}$ecules, or SIGMol, a non-autoregressive chemical language model based on continuous flow matching and flow maps that integrates naturally with training-free steering methods for molecular optimization. Compared to methods that require explicit post-training, we achieve competitive performance on a variety of small molecule property oracles $\textit{without any additional training}$, using only inference-time approaches based on both Feynman-Kac and gradient-based steering algorithms. The flexibility, performance, and demonstrated scaling of SIGMol suggest that it is well-suited as a generalist model for molecular design.
Chat is not available.
Successful Page Load