SpatialCell-JEPA: your neighbors describe you as well as you do, and you are in their picture too
Abstract
Understanding a cell in tissue means understanding it in its spatial context: in a tumor microenvironment, what a cell is doing depends on which cells surround it. Images give us appearance rather than function, so we test the measurable form of that premise, asking how much of a cell’s identity is recoverable from its neighbors alone. We train a predictor to estimate a cell’s frozen foundation-model embedding, together with a heteroscedastic noise term, from its spatial neighbor cells, and evaluate on patient-held-out cell typing across twelve Visium HD slides. The neighbor-derived estimate types as well as the cell’s own embedding. That embedding is withheld from the context set, but not from the neighbors’ image crops, which overlap it. The predictor’s attention shows why: it uses about seven of its thirty-six context cells, at a weighted mean distance inside the half-width of the crop the embedding was computed from. The result holds at a wider crop, and on a second foundation model. Spatial context is informative about cell identity, and the foundation model’s crop already spans the neighborhood the predictor reads, so reading it separately adds little. The equivalence holds only on average, and it conceals two opposite effects. Grouping cells by how many of their eight nearest neighbors carry the same type, the estimate is worse than the cell’s own embedding where no neighbor shares it and better where all of them do, turning from worse to better at about three of the eight. The estimate is therefore a reading of the neighborhood rather than a copy of the cell, and it is selective rather than uniform: it improves on the cell’s own embedding where the surrounding tissue is consistent with it, and the uncertainty it reports orders those regimes without using labels.