JouleChemRL: A Chemistry World Model-Powered Reinforcement Learning Framework for Oracle-Efficient Molecular Optimization
Abstract
In small molecule drug discovery, hit expansion and lead optimization require identifying molecules that simultaneously satisfy multiple, often competing, property objectives. Such multiparameter optimization (MPO) can be challenging because compound synthesis and experimental labeling remain time and resource intensive. Autoencoder-based generative models that perform continuous space optimization directly in a smooth and compact latent space have shown promise as a potential solution for efficient MPO, especially when paired with reinforcement learning (RL). Here, we introduce JouleChemRL, an autoencoder-based generative framework that extends the state-of-the-art proximal policy optimization (PPO) RL approach by incorporating a chemistry world model. We show that the world model, pretrained on the molecular property landscape in latent space and further fine-tuned online using oracle-evaluated molecules during RL, improves optimization performance over the standard PPO actor-critic framework. The world model framework also enables dream steps, where candidate molecules are evaluated directly by the world model in latent space rather than by the oracle, allowing additional model-based optimization steps without incurring the expense of oracle calls. Across a range of single and multi-objective molecular optimization tasks, JouleChemRL achieves comparable or superior optimization efficiency to state-of-the-art methods. Combining dream steps with periodic oracle evaluations substantially reduces the number of true oracle calls, anywhere between 20\% and 80\%, while maintaining comparable optimization performance. These results demonstrate that an adaptive chemistry world model can improve the oracle efficiency of latent-space RL for molecular optimization.