Which Mutations Matter When Calibrating a Protein Encoder Swap?
Adeliya Leleytner ⋅ Karina Romanova ⋅ Viliana Devbunova ⋅ Aleksandr Nikolich ⋅ Vladimir Platonov ⋅ Gregory Leleytner
Abstract
Protein foundation models are often used as frozen encoders beneath task-specific readouts. Replacing an encoder is cheap only if the old readout can be preserved, and one way to preserve it is to fit a map on paired, unlabeled sequences. We ask which sequences belong in that calibration set, and isolate one variable: whether their mutations occupy the residue positions used at deployment. In a controlled GFP crossover, two calibration sets have the same scaffold, size, mutation-order composition, and nearly identical pairwise identity. Across a fixed panel of 90 directed swaps among ten encoders, matched-position calibration improves assay Spearman by a median 0.067 over cross-position calibration, and 82 of the 90 swaps are positive. On a Nuclease B landscape the aggregate effect is 0.092 with 19 of 20 swaps positive, but one frozen primary side reverses, so the formal replication gate fails. The GFP effect exceeds a post-hoc intact-map random-arm null (one-sided Monte Carlo $p=0.01$). Mixed support is intermediate. Whether exact sequence pairing is needed remains unresolved. Calibration budget alone therefore does not characterize frozen-head transfer: unlabeled data should cover the mutation directions on which the head will be used. Position support is a useful diagnostic, not a certificate.
Chat is not available.
Successful Page Load