Towards Scalable Context-Aware Single-Cell Spatial Transcriptomics Prediction from Histology Images
Zijun Gao ⋅ Chunbin Gu ⋅ Xiangde Luo ⋅ Jinxi Xiang ⋅ Pheng-Ann Heng
Abstract
Predicting gene expression from H\&E-stained histology images offers a scalable alternative to costly spatial transcriptomics, yet most existing methods operate at the spot level, where signals from multiple cells are aggregated and critical cellular heterogeneity is obscured. Extending this paradigm to single-cell resolution is non-trivial. Naively applying pathology foundation models faces a scale mismatch: their patch-level representations mix multiple cells, whereas per-cell cropping or resizing distorts morphology and removes local context. Conversely, segmentation-based models without strong pretrained visual encoders often lack the morphological representation capacity needed for accurate molecular prediction and inherit errors from imperfect cell boundary masks. Here, we present CELLO, an efficient end-to-end framework that performs a single pathology foundation model forward pass per image and uses grid sampling to extract location-specific features for all cells simultaneously. We further introduce a distance-decay cross-attention module that refines each cell representation using spatially biased local morphological context. Using 52 paired Xenium--WSI samples spanning 12 organs and approximately 10M cells, CELLO achieves state-of-the-art performance in both in-distribution and out-of-distribution settings while delivering at least a 7$\times$ inference speedup over DeepSpot2Cell. Our work establishes a scalable foundation for single-cell gene expression prediction from H\&E images.
Chat is not available.
Successful Page Load