Training Large Language Models for Effective Small-Molecule Design with Synthetic Task Scaling
Abstract
Small-molecule drug discovery (SMDD) is a complicated optimization problem due to the combinatorial growth of chemical design space. Although frontier large language models (LLMs) have shown promise in navigating these vast and jagged design spaces, the limited control over model behavior and lack of customization for specific chemical domains motivate the need for open-weight alternatives. However, adapting open-weight LLMs to a drug discovery setting presents unique challenges because the evaluation setting is often computationally intractable for training, preventing the design of scalable post-training protocols around easily verifiable rewards. To circumvent this limitation, here we investigate the effectiveness of scaling post-training using synthetic data as a strategy for specializing open-weight LLMs to a drug discovery setting where the downstream reward is too expensive to optimize directly. By generating easily verifiable molecular design tasks based on low-cost cheminformatics oracles, we show how small open-weight models can be effectively post-trained using our scalable synthetic environments while generalizing to harder structure-based design tasks, reaching a level of performance comparable to frontier LLMs. We additionally show that by incorporating some harder environments using more expensive oracles, a curriculum training approach based on this hierarchy of synthetic design tasks improves performance, further bridging the fidelity gap between train and evaluation settings.