Pointwise regression matches or exceeds learning-to-rank strategies in ligand-based virtual screening
Abstract
Learning to rank (LTR) has been proposed as a natural fit for ligand-based virtual screening, where the goal is to enrich actives at the top of a ranked compound library rather than to predict activity values precisely. Previous studies report advantages of LTR models in virtual screening. However, these benchmarks primarily use target-based datasets with hit rates of 1-10% evaluated by normalized discounted cumulative gain (NDCG). Here we benchmark pairwise and listwise LTR against pointwise regression on 10 Gram-negative antibacterial phenotypic assays with hit rates of 0.06-1%, evaluating early enrichment across multiple molecular representations, model architectures, and split strategies. We find that pointwise regression matches or exceeds every ranking objective tested, and objective determines early enrichment more than representation or architecture. The gap narrows with increasing dataset heterogeneity. While LTR outperforms pointwise regression on a higher hit rate endpoint, the advantage collapses when hit rate is controlled.