HoldoutLab: A Four-Regime Benchmark of Evaluation Design for Gut-Microbial Drug-Metabolism Models
Abstract
The gut microbiome is increasingly recognized as a metabolic organ, capable of contributing to individual variability in drug response. The emergence of artificial intelligence models for microbial drug metabolism has been an active area of research, but many are evaluated on tasks that do not directly support realistic drug-development scenarios. We evaluate 21 published approaches and classify each by the drug-development task its evaluation supports: completing missing measurements among drugs and microbes already in training, stratifying an unseen community for a known drug, or predicting a structure the model has never seen. Evaluation concentrates on the first, while the claims describe the third. We take three approaches and evaluate them on all four tasks. Our results show that each approach performs worst on evaluating new chemical entities: a task that is central to drug discovery.