Provenance-Aware and Replica-Aware Evaluation of ML-Prioritized Cyclic Peptides against Staphylococcal α-Hemolysin
Abstract
ML-assisted molecular prioritization can be sensitive to seemingly incidental modeling choices as well as to stochasticity in downstream physical simulation. We examine both sources of uncertainty in cyclic-peptide prioritization against Staphylococcus aureus α-hemolysin (Hla). We first audit sequence-origin sensitivity in a protein-language-model pipeline by recomputing 827 archived 14-residue candidates across all cyclic rotations and comparing single-origin, canonicalized, and cyclic-permutation-averaged representations. The historical objective showed substantial ranking instability across sequence origins, with a median pairwise Spearman correlation of ρ = 0.690 and only 34% median top-50 intersection. Canonicalization and full orbit averaging both removed origin dependence yet remained materially different, sharing only 64% of their top 50; notably, the two candidates originally selected for molecular dynamics shifted from ranks 1 and 2 to 61 and 209 under full averaging. We then evaluated these candidates using three independently seeded 100-ns explicit-solvent MD replicas each. Candidate 2 showed more reproducible receptor association, with greater starting-contact retention and shorter initial-site distance in all 9/9 single-replica comparisons, while remaining receptor-associated in every analyzed frame and reproducing a cross-replica contact core. Together, these results show that both representation convention and trajectory seed can materially alter confidence in ML-assisted molecular prioritization, motivating explicit robustness audits before expensive experimental follow-up.