Knowledge graphs beat foundation models for gene perturbation prediction, but only the non-causal ones.
Abstract
Cheap gene-embedding methods not requiring large-scale single-cell gene perturba- tion datasets or the associated hardware are increasingly appealing and have proven competitive for gene perturbation prediction, at least on one public dataset. These embedding models rely on curated public knowledge graphs (KGs) that encode an inductive bias about what matters in describing a gene for downstream perturbation tasks. Whereas such embeddings were previously tested on a single KG, architec- ture, and unsupervised method, we benchmark their information content across multiple methods, disentangling the effect of the method from that of the KG itself. We find that non-causal, edge-based embeddings consistently outperform causal ones, foundation models, and baselines at predicting the perturbation response of significantly differentially expressed genes. Non-causal beating causal is the one pattern that holds across both regression and classification: causal embeddings trail the baselines on regression yet edge past them on classification, but never rival the non-causal ones.