Does Pretraining Teach Chemical Language Models Chemistry?
Abstract
Chemical language model (CLMs) are pretrained on large corpora of molecules with the aim of learning representations that transfer to diverse downstream tasks. Here, we systematically study a series of such CLMs across 64 molecular property prediction datasets covering classification, regression, activity cliffs, and virtual screenings. We find that supervised pretraining on molecular descriptors is consistently strongest, while much of the benefit of unsupervised pretraining can be recovered even after removing meaningful chemical structure from the pretraining data. Across all settings, classical fingerprints and descriptors remain substantially more accurate, robust, and efficient, and CLM embeddings provide no additional predictive benefit once these features are available. Our results suggest that CLM pretraining learns useful molecular representations, but surprisingly little chemistry beyond what classical molecular features already capture.