A Latent Ability Model for Benchmarking Molecular Representations
Nethaka Dassanayake ⋅ Andrew McNutt ⋅ Alexander S Rich
Abstract
The growing number of chemical foundation models calls for a standardized benchmarking method both for practitioners choosing representations and researchers identifying gaps. Previous work that ranked latent abilities of various chemical embeddings to predict ADMET properties suffered from significant shortcomings, including the use of biased datasets incongruent with real world usage and overstating statistically significant differences in model performance. We derive from first principles a variance inflation factor of $\frac{M+1}{3}$ for the Bayesian Bradley-Terry model evaluating $M$ models on $E$ endpoints at the exchangeable null and use numerical simulation to show variance inflation is still present in the non-exchangeable case. We propose a new method to directly model the margin of error as a function of latent ability and use it to evaluate models retrospectively on the datasets from the OpenADMET Polaris, ExpansionRx, and PXR competitions. We compare 34 embeddings and 14 end-to-end fine-tuned models across 17 ADMET endpoints and conclude that although no representation outperforms an ECFP baseline on a new endpoint at $p=0.05$, many neural representations outperform an ECFP baseline on most endpoints.
Chat is not available.
Successful Page Load