Post-Training LLMs for Property-Guided Molecular Optimization: Alignment Strategies and Multi-Metric Evaluation
Abstract
Molecular property optimization is a bottleneck in drug discovery: chemists iteratively modify structures to improve target properties while maintaining others, a slow process dependent on expert judgment. We present an end-to-end system that uses LLMs for property-guided molecular design, and compare three post-training paradigms --- supervised fine-tuning (SFT), SFT with a regression head, and Group Relative Policy Optimization (GRPO) --- across three model families and two endpoints: lipophilicity (logD) and inhibition of the hERG potassium channel. To evaluate outputs beyond accuracy we introduce a six-metric, chemistry-aware framework covering validity, synthesizability, property control, stability of secondary properties, novelty, and transformation size. Each metric is trivially satisfiable in isolation — a model returning its input unchanged is perfectly valid and stable — so the six are read jointly. Applied across ten base and post-trained models, the framework separates effects that a single score merges: in the general-purpose families SFT raises validity and synthesizability at the cost of novelty, GRPO recovers novelty while reducing the stability of secondary properties, and the regression head yields only a small, directionally consistent gain. No single metric orders the systems; the profile instead shows where each model sits among trade-offs a project has to choose between.