Critical Mass: Structure-Disjoint Splits Drift With the Libraries They Are Built On
Tomo Oga ⋅ Devesh Shah ⋅ Cailum M Stienstra ⋅ Gabriel Asher ⋅ Antonio Henrique de Oliveira Fonseca ⋅ Niall O'Connor ⋅ Michael Widrich
Abstract
Molecular machine learning evaluates out-of-distribution performance with structure disjoint splits: threshold a pairwise distance, then assign whole connected components to folds. For predicting structures from mass spectra, the field has converged on single-linkage clustering under the myopic Maximum Common Edge Subgraph (MCES) distance at a threshold of 10. This convention was set on 29,000 molecules but public libraries now hold 227,690 molecules. Because the MCES is NP-hard per pair over a quadratically growing set, the criterion has never been evaluated at that scale. We propose an optimized version of this criterion, allowing for computationally feasible application on the public libraries: admissible bounds and a union-find skip reduce $2.6 \times 10^{10}$ molecular pairs to only $1.15 \times 10^{7}$ integer programs, a $2{,}259\times$ saving, while provably producing the identical partition. We find that at this scale the graph percolates. A single component holds $95.5$% of the library, capping any test set at $4.5$%. What remains is not merely smaller, but systematically different: a held-out compound outweighs a training compound $88.1$% of the time. Ten edits is a large fraction of a small molecular graph and a small fraction of a large one, so heavy molecules are the ones that fail to connect. Large library benchmarks using this splitting rule naively therefore conflate structural novelty with extrapolation to larger, more complex chemistry.
Chat is not available.
Successful Page Load