Evaluating Negative Construction for Aptamer-Small Molecule Interaction Learning
Abstract
Negative examples are essential for training interaction models, yet experimentally confirmed non-interactions are scarce and are often replaced by synthetic negatives constructed under different biological and computational assumptions. We systematically evaluate six negative-construction strategies for aptamer--small molecule interaction prediction using fixed experimentally characterized test sets under stratified, aptamer-disjoint, and molecule-disjoint evaluation. Negative construction substantially changes downstream performance: under the same LightGBM pipeline, experimental-test ROC-AUC differs by more than 0.5 between synthetic strategies on the same split. Several methods achieve ROC-AUC above 0.95 on matched synthetic tests while dropping to 0.24--0.42 on experimentally confirmed non-binders under molecule shift, demonstrating substantial synthetic-to-experimental evaluation inflation. Experimental negatives provide the strongest supervision across all evaluated settings and predictive architectures. Moreover, using only 10\% of the available experimental negatives recovers most of the molecule-disjoint performance of the full experimental-negative set, while oversampling or synthetic expansion provides no additional benefit in this regime. The sensitivity to negative construction persists across four model families and fold-local leakage controls. These results show that negative provenance and experimentally grounded evaluation should be treated as explicit components of scientific dataset design rather than as incidental preprocessing choices.