Histology Reveals Proteomic Variation Beyond the Transcriptome
Mithil Shah ⋅ Akshath Sharma ⋅ Akhil Narayanan ⋅ Tanush S ⋅ Pranshu Nautiyal ⋅ Laksh Patel ⋅ Andrew Bae ⋅ Soham Batra ⋅ Siddharth Karuturi ⋅ Oliver Loy ⋅ Aarav Lala ⋅ Patrick Feng
Abstract
Molecularly grounded pathology foundation models connect routine H\&E images with tumour biology, but most are supervised using RNA rather than protein abundance. This is an important limitation because RNA does not fully represent the proteome: across 715 tumours from eight CPTAC cohorts, a protein's own transcript explains a median of only 17.2\% of its within-cohort variance. We hypothesize that H\&E morphology captures protein-level variation not already explained by RNA. We extract UNI2-h patch features from 1,692 tumour whole-slide images, aggregate them by gated attention multiple-instance learning, and predict mass-spectrometry abundance for 13,509 proteins under patient-level cross-validation, with targets centred within cancer type so prediction cannot be driven by tissue identity. After removing the variation explained by each protein's own transcript, morphology still predicts the remainder at $r=0.233$ (SD $0.025$) against a within-cohort permutation baseline of $r=0.032$; raw protein prediction reaches $r=0.276$, so removing transcript-associated variation costs little. Comparing four ridge models under identical splits---cognate transcript ($r=0.414$), full transcriptome ($r=0.384$), morphology alone ($r=0.285$), and transcriptome plus morphology ($r=0.466$)---adding morphology improves prediction by $\Delta r=+0.082$ (fold-level SEM $0.007$), with 84.6\% of 11,997 proteins improving and 62.8\% improving in all three folds against a 12.5\% chance expectation ($p<10^{-300}$). Gains are largest among the most reliably measured proteins, making measurement noise an unlikely explanation. Routine H\&E therefore contains proteomic information that RNA does not capture.
Chat is not available.
Successful Page Load