When perturbation queries fail: static gene and protein embeddings outperform simulated scFM perturbation responses
Abstract
Regulatory structure has been shown to be extracted from pretrained single-cell foundation models (scFM) by querying how its output changes when one gene’s expression is computationally altered. We ask whether such queries carry information about the response actually measured when that gene is knocked down. On two genome-scale CRISPRi Perturb-seq screens (HepG2 and Jurkat) we extract such features from three released scFMs and read them against the same models’ static gene embedding tables and against baselines that use no foundation model at all. Every feature set is fed into a simple MLP translator, predicting only the perturbation-specific residual, scored on held-out groups of similar-responding perturbations. The static tables and raw ESM2 protein embeddings substantially outperform the query-derived features, which sit near co-expression features, and fine-tuning scGPT on the perturbation data does not change this. Read from hidden states instead of the output head, the same queries show the perturbed gene’s iden14 tity is decodable at every depth and beat their output-head counterparts, yet still fall short of a static table. Our results show that strong performance on regulatory16 edge benchmarks does not, by itself, establish that the same query predicts the consequences of a perturbation.