Beyond the Final Layer: Evaluating Single-Cell Foundation Model Representations for Molecular Prediction
Abstract
Single-cell foundation models (scFMs) are often benchmarked using final-layer embeddings, but layer choice is rarely tested. We evaluate Geneformer, scBERT, scGPT, and scFoundation as frozen feature extractors across cell type prediction and perturbation-response prediction, combining layer-wise probes with leave-one-dataset-out (LODO) selection, readout controls, transfer analysis, representation diagnostics, centered kernel alignment (CKA), and perturbation retrieval/collapse checks. Final-layer-only evaluation can miss useful intermediate representations, although no single non-final layer is universally optimal. In LODO-controlled cell type prediction, scGPT layer 2 and scBERT intermediate layers significantly improve paired held-out accuracy on multiple datasets after Benjamini--Hochberg (BH) correction. In perturbation prediction, non-final layers often match or exceed final-layer performance, but optima vary across datasets and metrics. These results support reporting extraction depth, readout, and perturbation diagnostics as standard components of scFM benchmarking, and suggest that when non-final layers are on par with final layers, early extraction can serve as an early-stopping strategy that avoids feed-forward computation through the remaining layers and reduces unnecessary computational cost.