GeneSpeak-FP: Target and Compound Retrieval from Observed Cell-Level Perturbation Signatures
Abstract
Single-cell perturbation profiles may enable retrospective identification of known compounds across cellular contexts. We present GeneSpeak-FP, a gene-token Transformer that jointly ranks recorded compounds and compound-associated target annotations from a treated cell relative to a cell-line-specific DMSO reference, together with a coarse organ token. On Tahoe-100M, we derive a within-compound 90/10 split from 11,673 pre-sampling drug-cell-line groups, with no complete group shared across partitions, although compound and cell-line identities may recur individually. Across a 38,400-query Monte Carlo pass, GeneSpeak-FP achieves compound Hit@1/10 of 0.129/0.343 and target Recall@10/20 of 0.408/0.544 over fixed banks of 379 compounds and 278 target genes. At K=10, these values are approximately 13x and 11x their analytic chance levels; compound Hit@1 is nearly 49x chance, and compound Hit@10 exceeds the strongest evaluated post-hoc bag-of-genes control by more than 9x. A separately executed stochastic pass from the same validation pool yields nearly identical aggregate values. These results establish a large and consistent closed-library retrieval gap for held-out drug-cell-line contexts. Because compound identities may recur and target labels are non-exhaustive metadata, the result supports retrospective candidate-space narrowing rather than unseen-compound discovery, causal target identification, or clinical use.