IDPInterp: Towards Interpretable Intrinsic Disorder in Protein Language Models
Trevor Xing-Xie ⋅ Daniel Moyer
Abstract
Intrinsically disordered proteins are responsible for myriad gene expression pathways and disease mechanisms, including key implications in neurodegenerative diseases and hematological cancers. Despite composing an estimated 30-40\% of the human proteome, the application of novel computational methods to intrinsic disorder, such as interpretability, is limited. Protein language models (PLMs) encode intrinsic disorder despite training only on sequence, yet what they represent about disorder and how their training objective shapes it is not known. We train sparse autoencoders (SAEs) on PLMs with contrasting training objectives, evolutionary covariation (ESM-2) and biophysical dynamics (SeqDance), and interpret the resulting disorder features with region-aware evaluation and held-out-validated automated descriptions. In ESM-2, disorder is represented as a compositional basis of roughly 45 monosemantic feature families organized by amino acid composition and charge pattern, while conformational sub-states such as the molten globule and partner-induced folding are not linked with any individual feature. Comparing objectives, both models encode charge-driven compaction, but only the dynamics-trained model encodes hydrophobic collapse, the driving biophysical mechanism of the molten globule. Linear probes find stronger molten globule detection in the weakest SeqDance layer than the strongest of twelve ESM-2 layers ($+0.225$ at matched architecture, scale, and depth, $p = 0.0007$), yet neither training objective gives uniformly stronger disorder representations.
Chat is not available.
Successful Page Load