OligoBench: Benchmarking Foundation Models for Biological Feature Discovery in Oligonucleotide Design
Abstract
Machine learning approaches to oligonucleotide design typically predict activity from predefined molecular representations, leaving much of the biological context of the RNA target outside the representation-learning problem. We introduce OligoBench, a framework for benchmarking whether general-purpose foundation models can autonomously discover biologically informative features for oligonucleotide design. Rather than predicting activity directly, models receive an oligonucleotide and its complete target transcript, without the binding site or activity label, and return executable feature functions evaluated through fixed downstream predictors. Across four activity datasets spanning antisense oligonucleotides (ASOs) and siRNAs, foundation models from multiple model families consistently improve over oligonucleotide-only baselines. The models recover relevant target context, construct distinct representations, and produce clear differences in predictive gains across model families and datasets. OligoBench provides a reusable benchmarking framework for comparing biological feature-discovery capabilities across foundation models and for extending evaluation to new model generations, RNA-targeting modalities, datasets, and prediction endpoints. Fully reproducible code for running the experiments, and for evaluating any new model in the same harness, will be available at https://anonymous/oligobench.