From Performance to Decision: Evaluation, Not Architecture, Drives Antimalarial Virtual-Screening Conclusions
Marvellous Ajala
Abstract
Machine-learning virtual screening is routinely judged by accuracy and ROC-AUC on randomly split data, producing two recurring narratives: that deep learning outperforms simple baselines, and that severe class imbalance makes screening models useless. We investigate both narratives rigorously for antimalarial activity prediction on the ChEMBL Legacy dataset (65,751 train / 112,994 test compounds) under a strict structural split (Lo-Hi, Tanimoto dissimilarity $\geq 0.4$). We compare seven classical classifiers over two molecular representations with a pretrained chemical transformer (ChemBERTa-2) and three graph neural networks (GCN, GAT, MPNN), using multi-seed repeats, bootstrap confidence intervals, DeLong and paired-bootstrap tests, McNemar tests, class-weighted training, enrichment metrics (EF, BEDROC, recall@k), simulated screening campaigns, imbalance sweeps, learning curves, and a random-vs-structural split comparison. Three findings emerge. First, neither simple nor deep-learning models dominate: descriptor-based Random Forest, MLP, and XGBoost (AUC $\approx 0.68$) and the pretrained transformer ChemBERTa (0.683) form a statistically indistinguishable top cluster, while the naive GNNs are significantly worse (GAT 0.658, MPNN 0.594, GCN 0.586). Second, class weighting raises recall $3$--$50\times$ with no ROC-AUC loss in every model family, showing that the apparent ``catastrophic recall'' reflects default-threshold, unweighted evaluation rather than model failure. Third, ranking translates into usable screening value: top-1\% enrichment reaches $5.5\times$, and a simulated screen recovers 5\% of all actives by testing $\sim$1,000 compounds at $\sim$96\% top-1k purity. Accuracy is uninformative (all models sit at the 0.827 majority floor) and random-split over-optimism is modest ($0.01$--$0.05$ AUC). The failure mode is the evaluation, not the models.
Chat is not available.
Successful Page Load