HUME: High-Throughput Molecular Embeddings from Fingerprints and Descriptors
Abstract
Learned molecular embeddings promise to replace hand-crafted representations with compact, transferable features, yet classical fingerprints and descriptors remain remarkably difficult to beat. We argue that fingerprints already capture molecular structure with near-Weisfeiler–Lehman resolution, while descriptors supply the smooth physico-chemical information we seek from learned embeddings. We introduce HUME, a high-throughput embedding that combines both, computing fingerprints and a comprehensive set of exact descriptors in a unified optimized implementation. Across 26 scaffold-split property-prediction datasets, HUME outperforms state-of-the-art learned embeddings, which provide no additional predictive benefit when added to HUME. Crucially, HUME is also faster to compute than learned embeddings, undercutting the traditional scalability argument for learning these features.