Optimizing Fusion of Foundation-Model Gene Expression Programs for Verification against Biological Evidence
Abstract
Foundation models (FMs) trained on transcriptomic data can propose candidate disease-related gene expression programs (GEPs) directly from their learned embeddings, without checking these programs against curated biological knowledge at generation time. One possible solution is to treat agreement on GEPs, generated by independently trained FMs, as an annotation-free signal to judge whether the resulting programs are trustworthy. We test this idea in a case study on endometriosis transcriptomic data. Harmonising embeddings from two architecturally distinct RNA FMs, Geneformer and BMFM-RNA, via Grassmannian manifold geometry, we derive ten consensus GEPs: five show strong cross-model agreement and five are architecture-specific. We then validate each GEP against two independent verifiers built from the same cohort and attribution pipeline: curated Gene Ontology (GO) annotation, and a gene panel constructed via a combinatorial optimisation (CO) algorithm, Weighted Maximum Independent Set (MIS), over a co-expression graph. In this case study, the two verifiers disagree. The co-expression gene panel shows substantially higher gene overlap with high-agreement GEPs than with low-agreement ones (42--58\% vs. 4--20\%), while GO annotation confirms only a scattered subset, with no clean split by agreement level. We further show that when we apply a second CO algorithm (Set Packing, maximising attribution-weighted coverage subject to non-overlap) to prune redundant GEPs from the hypergraph, it excludes two high-agreement GEPs that GO annotation would have confirmed. Cross-model agreement is therefore informative but not a substitute for validation as it correlates with one biologically grounded verifier but not another, so it cannot substitute for a biological-validity check on the redundancy-elimination output. The Set Packing step optimises attribution-weighted coverage alone and biological validity is checked only afterward as a separate filter. This motivates further work on making redundancy-elimination for FM-generated hypotheses biology-aware from the outset. Building curated biological knowledge into the optimisation itself, as a constraint rather than a post hoc filter, would let evidence strength and biological validity be optimised jointly instead of sequentially.