Evaluating scConceptX, a Multi-Species Transcriptomics Foundation Model, by Agentic Rediscovery of Known Biology
Mojtaba Bahrami ⋅ Fabian Theis
Abstract
Single-cell foundation models are trained on ever larger and more diverse data, yet their evaluation remains confined to a limited set of benchmarks that are hard to scale. At the same time, the literature holds a vast body of published biological findings that could serve as tests, but turning each one into an executable evaluation has required prohibitive manual effort. We introduce scConceptX, an extended version of the scConcept contrastive model, scaled to 200M parameters, full-transcriptome inputs and 360 million cells from 16 species, with gene tokens that combine frozen protein-language-model embeddings with species-specific learnable residuals. To evaluate it, we propose agentic literature recovery: an LLM agent surveys the literature, selects studies, extracts quantitatively testable findings, locates and harmonises the public data, designs tests with pre-declared pass rules, runs the model, and audits its own conclusions. With a handful of short prompts, the agent re-tested 30 published findings from eleven studies (five single-species studies in human, mouse, marmoset, zebrafish and pig, and six cross-species studies spanning 10 species from human to zebrafish), using the pretrained model and scConceptX+, the same model adapted self-supervisedly on each study's own cells. Within single species, adaptation raised recovery from 6/13 to 10/13 findings. Across species, zero-shot scConceptX recovered 8 of 10 findings about cell identity, homology and shared cell states, including 57 consensus primate cell types, retinal classes across 430 million years, and the inflammatory fibroblast state shared by human psoriasis and a mouse model of epidermolysis bullosa; adaptation sharpened cross-species cell-type mapping (57-type balanced accuracy 0.65$\to$0.75) but eroded two findings that depend on species-specific differences. Because the agent carries out every step, from finding the studies to auditing the tests, this kind of hypothesis-level evaluation of foundation models against the published literature becomes practical at scale.
Chat is not available.
Successful Page Load