Heterophily in Spatial Omics: Understanding Graph Convolution Failure and Recovery
Abstract
Graph neural networks are widely used to model cell-cell relationships in spatial omics data, but standard graph convolution assumes that neighboring cells tend to have the same labels. We examine how well this assumption holds across tissues and how violations of it affect supervised cell-type prediction. Across five public datasets using antibody imaging, single-molecule FISH, and spot-based transcriptomics, adjusted homophily ranged from 0.17 in hypothalamus to 0.85 in human cortex. Lower-homophily tissues had more cells at cell-type boundaries, where neighboring cells often had different labels. Model performance followed the same pattern. In breast carcinoma, GCN accuracy drops to 0.269 on interface cells, compared with 0.714 for a graph-free MLP. In cortex, the same architecture benefits substantially from graph information. We then test whether learning how neighboring cells influence each other improves performance under heterophily by introducing SheafST, a cellular sheaf diffusion network with learned d x d restriction maps. However, the learned maps provide little additional benefit, with accuracy changing by only +0.004 [-0.017, +0.025] when they are replaced by identity maps. Instead, the most important factor is whether the model maintains a separate path for each cell's own representation. Adding this path increases GCN interface accuracy from 0.269 to 0.709, a paired gain of +0.440 [+0.408, +0.471], bringing its performance in line with GraphSAGE, H2GCN, and SheafST. With this added path, graph-based models do not consistently outperform the graph-free MLP on interface cells across the four heterophilous datasets. In cortex, graph information improves performance by +0.142 [+0.098, +0.187]. These results show that preserving each cell's own representation is key to GCN performance under heterophily, but adding spatial information does not consistently improve prediction over a graph-free MLP.