The Effect of Preprocessing Choices on Single-Cell Foundation Models
Abstract
Between the FASTQ read file generated by single-cell RNA-sequencing (scRNA-seq) experiments and the count matrix consumed by a foundation model is a series of preprocessing steps. These steps include count matrix denoising, count normalization, highly variable gene selection, value encoding, and neighbor-graph construction. While the parameters of each step are often chosen arbitrarily, they can have a large effect on the output of a model. Four popular scRNA-seq foundation models — scGPT, Geneformer, scFoundation, and scVI — were fine-tuned or trained with different combinations of these preprocessing steps. On average, changing any one step replaces 86% of a cell's nearest neighbors and leaves the cell's embedding at a cosine similarity of 0.85 to the default. When evaluated on cell-type prediction, changing any one step relabels up to 9% of held-out cells, and up to 20% of cells are sensitive to at least one preprocessing step. Stacking several changes into a complete alternative pipeline relabels no more cells than the largest single change. The relabeled cells are overrepresented in rare populations that a different step can erase entirely. In some datasets and models, changing normalization alone decreases accuracy by up to 13 percentage points. If the preprocessing recipe differs between fine-tuning and inference, accuracy can drop to levels of random chance. These results demonstrate the importance of preprocessing choices in scRNA-seq foundation models, and we provide recommendations for making such choices.