FidelityShift: Data-Readiness Lessons for Protein-FM Evaluation Under Alternate Decoding
Abstract
Protein foundation models are increasingly used as transferable priors for biological prediction, yet disease-state proteomes can contain products not directly encoded by canonical sequence. We turn established tryptophan-to-phenylalanine (W→F) alternate-decoding immunopeptidomics into a verification benchmark: given W→F events observed in training dataset-defined context families, rank candidate altered events for a held-out context without constructing false negatives. Across 224 eligible events, 18 dataset-defined context families and 18 leave-one-context-family-out folds, the tested static sequence-only frozen ESM2, ProtT5 and ESM2 zero-shot scores are far weaker than event-history recurrence priors. Exact-HLA pMHC scoring is stronger than tested static sequence-only frozen FM scores on the 13 of 18 exact-HLA-supported folds, but it also does not close the gap. W0 event-history recurrence exceeds parent-protein recurrence in all 18 folds (mean Recall@10 delta 0.0774, 95% CI [0.0695, 0.0848]), while within-parent evidence is positive but limited. FidelityShift is therefore a verification-driven behavioural benchmark and information-source diagnostic: the tested static sequence-only frozen FM scores do not recover the dominant event-history recurrence structure, without implying that all protein FMs or future state-aware systems must fail.