Slide-level batch structure limits histology-guided supervision of transcriptomic foundation models
Abstract
Spatial transcriptomics pairs spatially resolved gene expression with tissue morphology in the same tissue section. Transcriptomic foundation models provide general-purpose gene-expression representations, but these representations can retain slide- and cohort-specific variation that obscures biological signal. Here, we test whether matched H&E histology can improve these representations as a training-time supervisory signal that is discarded at inference. Across three transcriptomic foundation models, histology-guided supervision provides no consistent aggregate improvement in cross-donor annotation transfer, despite substantially stronger transfer from histology alone. We find that gene-expression embeddings from spot-based spatial transcriptomics are low-dimensional and strongly structured by slide identity, limiting the shared gene--morphology signal available for cross-modal transfer. Guidance benefits some morphologically distinctive classes but reduces held-out gene predictivity broadly across the transcriptome. These results suggest that cross-modal supervision can reorganize information already encoded in a frozen representation but cannot recover information the representation does not retain, highlighting the importance of diagnosing slide-specific structure before applying such supervision. Code availability: https://anonymous.4open.science/r/vision-guided-transcriptomics-fm-2026-75FE