From ranking to prediction: Label Construction Shapes Aptamer–Protein Binding Models
Armin Geraili ⋅ Arman Seyed-Ahmadi ⋅ Hossein Zargartalebi ⋅ Dingran Chang ⋅ Shana O Kelley
Abstract
High-throughput selection platforms can assay millions of aptamer candidates against a protein target in a single experiment, but they produce sequencing counts rather than direct binding labels. Converting these counts into supervised learning targets is therefore a critical modelling choice whose impact is rarely studied. Here, we report the first predictive aptamer–protein binding models trained on ProSELEX data from an experiment targeting myeloperoxidase (MPO). We compare two approaches for constructing binding labels: binary presence/absence and enrichment-based labels that incorporate information from the sorting process. Using presence/absence labels, we benchmark 109 configurations spanning k-mer classifiers, DNA/RNA foundation models, SELEX-pretrained transformers, and existing aptamer-specific predictors. We then retrain six representative models with enrichment-based labels while keeping all other settings fixed. With presence/absence labels, no model exceeds an AUROC of 0.63, whereas enrichment-based labels improve AUROC by 0.20–0.23 across all six models. Masked-language pretraining on ProSELEX sequences provides a further improvement of 0.133, yielding a best AUROC of 0.860, with sensitivity of 0.773 and specificity of 0.836. We further evaluate the models blindly on two independent experimental panels: 38 published aptamers with flow-cytometry-measured $K_d$ values and 21 aptamers measured by biolayer interferometry. The in-distribution performance does not transfer to these panels: measured non-binders fall within an intermediate enrichment region excluded during label construction, leaving the models without experimentally validated negatives in their learned negative region. These results show that label construction can influence predictive performance more strongly than model architecture while also defining the limits of out-of-distribution generalization.
Chat is not available.
Successful Page Load