Benchmarks Design under Data Scarcity: From Coarse Labels to Diagnostic Evaluation of Biosynthetic Gene Cluster Models
Abstract
Biosynthetic gene clusters (BGCs) are co-located genes that encode the biosynthesis of secondary metabolites, a major source of clinically used antibiotics and anticancer agents. Recent advances in deep learning, particularly self-supervised foundation models, have spurred growing interest in BGC sequence modeling, but evaluation infrastructure has lagged behind. Current practice often relies on coarse classification accuracy over small, experimentally curated datasets, making it difficult to discriminate model capabilities or assess downstream utility. We introduce BGC-Bench, an evaluation suite designed to resolve this limitation by systematically incorporating biochemical knowledge in the benchmark design. In terms of generalization, we test whether representations transfer to genetically and chemically novel samples; for functionality, we examine application-relevant tasks, including retrieval, prediction, and active learning. BGC-Bench reveals that no single model class dominates: pretrained foundation models perform strongly under out-of-domain generalization but remain sensitive to training conditions; compositional baselines are competitive and lead in cross-modal retrieval; and active-learning gains vary substantially across tasks, strategies, and model families. Beyond these BGC-specific findings, BGC-Bench highlights a broader principle: domain-informed benchmark design over existing datasets enables better evaluation in data-scarce scientific domains.