Toxicity Prediction Tools Do Not Generalize to Novel Chemistry
Martin Weiss ⋅ Circe Hsu ⋅ Antonio Henrique de Oliveira Fonseca
Abstract
[Tiny paper submission — 4 pages.] Computational toxicity prediction rests on an assumption that is rarely tested directly: that molecular structure carries enough information to anticipate whether a compound will have adverse effects in human patients. We benchmark toxicity prediction methods on a curated panel of $569$ approved drugs labeled by regulatory outcome, each below $0.70$ ECFP4 similarity to every compound in invitroDB, and evaluate $35$ models spanning classical QSAR, fine-tuned foundation encoders, and toxicity-targeted tools. No learned representation outperforms a plain fingerprint, suggesting the bottleneck is not the sophistication of the representation. Of the tools trained specifically for toxicity prediction, only one performs above chance, and none matches what approval year alone explains, a predictor containing no chemistry at all. No tool dedicated to drug-induced liver injury (DILI) reliably beats chance on our panel, and the external set where some succeed is one that approval year alone largely predicts. These results point to a limit of molecular structure as an input, and suggest that progress will depend on inputs that capture exposure and metabolism rather than better models over structure alone.
Chat is not available.
Successful Page Load