Downstream Transfer of Scaling in Molecular Language Models
Abstract
Pretraining loss of language models scales predictably. Downstream performance often does not. We test that gap for chemical language models trained on PubChem-derived molecular data, using four sizes of networks (170M, 380M, 1.3B, 3B) and constant data size. We evaluate in three downstream application categories: zero-shot molecular generation conditioned on basic properties, docking-based molecular optimization, and property prediction with fine-tuning. Generation improves with scale, even when conditioned on three properties. Molecular optimization scaling, as measured on a recently proposed PMO-Dock benchmark, is inconsistent: on average larger models generate better molecules for two out of three benchmark categories, while the trend is non-monotonic for certain protein targets. Fine-tuning results are also mixed: a weak monotonic trend on Polaris ADME suite of tasks, inverted-U curves on PXR induction and BELKA. The results indicate that training much larger models might not be justified for many downstream applications, and motivate future research on more scalable fine-tuning and optimization recipes.