scXplain: Dissecting single-cell foundation models beyond predictive performance
Abstract
Single-cell foundation models (scFMs) are increasingly used for in-silico biological discovery, yet whether their internal computations align with known molecular mechanisms remains poorly understood. We systematically investigate the learned representations and decision-making processes of two scFM families, Cell2Sentence and scGPT, across blood and brain cell annotation tasks. We employ attribution methods to assess whether model predictions rely on established marker genes, revealing substantial marker recovery alongside architecture-dependent positional biases and distinct failure modes. Using symbolic explanations, we additionally show that the relevance of markers shared by closely related cell subtypes depends on the broader gene context, providing evidence of higher-order, context-sensitive gene interactions. Together, these results demonstrate the potential of attribution-based explainability techniques to characterise the conformity of scFM predictions with biological priors, enabling finer insights for investigating generalisation and for model debugging.