Match and Correct: A Repeated Computational Motif in Evoformer Layers
Ishan Khire ⋅ Naren Manikandan ⋅ Tyler Hayes ⋅ Giri Krishnan
Abstract
Triangle attention and triangle multiplicative updates form key components of the Evoformer, and were introduced specifically to encourage consistency with the triangle inequality. Yet it has not been tested whether representational changes over layers reduce triangle inequality violations. We probe the pair representation at each of OpenFold's 48 Evoformer layers, decoding final C$_\alpha$--C$_\alpha$ distances with four probe families (a shared linear probe, per-protein linear probes, a sequence-separation-gated linear probe, and XGBoost) and reconstructing intermediate 3D structures from the decoded distances via multidimensional scaling. Across a cohort of 552 CAMEO and ECOD proteins, distance decodability ($R^2$) and reconstruction quality (TM-score) both rise in early layers, plateau in the middle of the network, and rise again in late layers. Directly measuring triangle-inequality violations in the decoded distances reveals a different, non-monotonic pattern: violations rise to a peak, fall to a trough, rise to a second peak, and only then decline, tracing out two distinct episodes of geometric refinement rather than one continuous improvement. Cumulative forward and reverse layer-wise ablations of the Pair Transition, MSA Transition, and triangle mechanisms localize each episode to different Evoformer components, suggesting that the two stages of geometric reasoning coincide with the use of the Pair Transition MLP and the triangle mechanisms. These results demonstrate how model interpretability leads to candidate mechanistic principles that can help identify the unifying laws of structure determination of proteins.
Chat is not available.
Successful Page Load