Quantifying the value of spatial data for pretraining by using a multi-modal model integrating histology, bulk-RNA and spatial transcriptomics
Abstract
Spatial Transcriptomics (ST) is a method that provides transcriptomics data associated with positional information within the tissue, it has contributed greatly to our understanding of the key role the spatial organization of cells and tissues have in the progression of disease. However, this technique remain expensive, which may limit its adoption in clinical pipelines in the near future. We present a multi-modal masked model that integrates histopathology (H&E), bulk-RNA sequencing and spatial transcriptomics data into a joint representation that can be used for downstream prediction tasks. We show that when inference is restricted to H&E data, models trained with H&E and ST (with or without bulk RNA) significantly under-perform models trained using only H&E and bulk RNA, even after controlling for differences in dataset size. These findings highlight the need for benchmarks that explicitly evaluate spatial tasks. They also suggest that bulk RNA sequencing should serve as a control when assessing improvements in representation quality and downstream-task performance attributed to ST data.