Measuring What Histology Can Predict: Leakage-Stratified Evaluation for Spatial Transcriptomics
Sajib Acharjee Dip ⋅ Liqing Zhang
Abstract
Models that predict spatial gene expression from H\&E histology are judged by one number: the per-gene correlation between predicted and measured expression. We show that number can move by an order of magnitude without touching the model. On the $360$ slides and $830{,}000$ spots of the $12$ analysed organs, holding the data, the genes and the predictor fixed, three choices in how the evaluation is built decide the result. The first is what you hold out. Moving from random spot splits to holding out whole studies costs a factor of $1.3$ to $6.7$. No step ever raises accuracy, and the spot-to-study drop holds in all $12$ organs. Patient-level splits, the current standard, buy almost nothing \emph{here}, because most slides in this corpus carry a single donor; the rung is worth exactly what a corpus's donor-per-slide ratio allows. The second is whether correlations are pooled across slides: under pooling, a predictor that is told each slide's average expression and nothing else recovers $59$--$99\%$ of the reported score ($59$--$87\%$ if the two-study organs are set aside). The third is the missing denominator. We estimate how reliably each gene is measured in the first place, and we also measure that estimator's floor, which turns out to be high enough to matter. Holding out studies also shrinks the training set, so we regroup slides at random into fake studies of the same sizes. Between $58\%$ and $77\%$ of the drop survives, which shows it comes from the grouping and not from having less data. The same pattern appears for four kinds of predictor and three image encoders, including a ResNet-50 that has never seen a tissue slide. We release the protocol as tested software.
Chat is not available.
Successful Page Load