Cell-level representations in pathology foundation models are a question of read-out
Abstract
Pathology foundation models are pre-trained and deployed at the tile or slide level, yet many downstream tasks are defined at the cell level. The common practice is to crop a region around each cell, resize it to the model’s input size, and extract the CLS token. This forces a trade-off between cell specificity and context preservation: crops tight enough for the cell to dominate the field of view lose the cell's context, while crops at native scale leave the cell occupying a small fraction of the tile, lacking a special focus on the cell itself. An alternative is available at no additional cost. Recent pathology foundation models are trained with DINOv2, whose iBOT objective explicitly supervises token-level representations, so the patch tokens of a single tile already form a grid of local descriptors. We can therefore obtain cell embeddings by mapping each cell centroid to its corresponding token(s) (\textit{spatial indexing}). We compare the two read-outs across three pathology foundation models (H-Optimus-1, UNI2, Virchow2) and three Xenium datasets that we compiled and annotated. For cell type classification, spatial indexing consistently outperforms crop-resize and exceeds the PanNuke trained CellViT baseline. For gene expression prediction with DeepSpot, it improves Pearson correlation while making the neighborhood-aggregation module unnecessary. These results indicate that existing pathology foundation models already encode transferable single-cell representations, and that accessing them is a question of read-out rather than retraining.