Multi-Oracle Agreement Reveals the Limits of Self-Consistency Evaluation in RNA Design
Abstract
Computational RNA design methods routinely evaluate designed sequences by refolding them with the same structure-prediction model used during optimization, measuring agreement with the design target. While scalable, this practice carries a fundamental risk: when optimization and evaluation share the same underlying model, high scores can reflect model-specific idiosyncrasies rather than genuine structural quality. To characterize the severity and mechanism of this problem, we audited nine RNA design methods spanning four algorithmic paradigms, evaluating approximately 11,000 sequences against predictors from two independent model families: physics-based thermodynamic models and machine-learning models, alongside five tertiary-structure predictors and large-scale SHAPE chemical-mapping data. We find that 78.1% of designs produced by single-objective optimization pass one predictor yet fail another, while ensemble-based methods reduce this disagreement 30-fold. A controlled experiment identifies optimization objective breadth, not search strategy, as the primary driver of this fragility. Cross-predictor disagreement carries a calibrated experimental signal: lower disagreement predicts higher SHAPE accuracy across 40,000 sequences, yet remains null-to-anti-predictive for biological function, establishing a clear scope boundary. We release ACCORD (Across-oracle Calibration for COmputational RNA Design), a benchmarking workflow that reports per-predictor scores, experimentally calibrated uncertainty tiers, and explicit warnings for invalid use cases, providing the community with a principled foundation for RNA design evaluation.