Calibrated Uncertainty Quantification on TDC ADMET: What Can Be Resolved Across 22 Datasets?
Abstract
Absorption, distribution, metabolism, excretion, and toxicity (ADMET) models often rely on small, heterogeneous datasets, yet their predictions may determine which compounds proceed to experimental testing. Useful evaluations must therefore assess both predictive performance and the quality of reported probabilities or intervals. We ask which common uncertainty-quantification methods perform best on the 22 Therapeutics Data Commons (TDC) ADMET datasets and what this benchmark can resolve. We compare eleven calibrated pipelines using nested five-fold cross-validation across ten chemical partitions and five model seeds. Ensemble Monte Carlo (MC) dropout has the best average rank for TDC performance, Brier score, and 90% interval score. For TDC performance, however, it is not separated from the other ensemble pipelines; for Brier and interval score, LightGBM is also not ruled out as best. LightGBM wins more individual datasets than any other method for TDC performance and Brier score, but performs poorly on several endpoints. Smaller evaluations can identify a different leader and often fail to recover the broader comparison. These results make LightGBM and an ensemble neural model complementary baselines. Comparative claims should be checked across multiple chemical partitions using statistical analyses matched to the question being asked.