Darwin-7B: A Multi-Omic Foundation Model for the Human Gut Microbiome via Sparsified Quality-Aware Tokenization
Abstract
Microbial ecosystems govern human health, agricultural productivity, and biogeochemical cycling. Two coupled modalities determine their function: which organisms are present (metagenomic content) and which small molecules they produce and consume (metabolomic content). Existing biological foundation models cannot represent this joint state: they operate on a single modality, on assembled single-organism genomes, and on the within-DNA axis of generalization, leaving the multi-omic, mixed-organism, mixed-quality regime that governs clinical and surveillance use unaddressed. We present Darwin-7B, a 7-billion-parameter multi-omic foundation model whose primary axis of generalization spans both modalities (metagenome + metabolome) and organizational scales (read → community → host phenotype). We combine a Mamba–Transformer hybrid backbone (24 of 32 layers Mamba) with a hypergraph neural network over KEGG metabolic-reaction hyperedges and bidirectional cross-modal attention modules that align genomic and metabolomic representations in a shared 40,192-token embedding; an Aitchison-space compositional consistency loss respects the simplex geometry of microbial-abundance data at the output. We pretrain on 8 T base pairs of metagenomics, 250K LC-MS/MS metabolite profiles, and 2M KEGG/GO functional annotations, preparing the metagenomic corpus with a published quality-aware tokenization framework. Against the matched 7B genomic baseline METAGENE-1, we report 94.5 MCC on Pathogen Detection (vs. 93.0) and 0.98 F1 on CAMI species-level metagenomic profiling; on a December 2025 SRA outbreak benchmark post-dating the pretraining cutoff, we report 91.0 MCC vs. 81.2. We unlock four clinical tasks single-modality genomic models cannot reach: IBD AUC 0.947, T2D AUC 0.883, antibiotic-resistance AUC 0.910, and metabolic-pathway prediction wF1 0.91, beating the strongest multi-omic baseline (MOGONet) by +3.0–+3.7 AUC. Token-matched scaling from 100M to 70B yields monotone gains with no saturation on the multi-omic clinical AUCs. We further validate predictions in a prospective wet-lab pilot (n = 87) at predicted-vs-observed Pearson r = 0.72. Sparse-autoencoder probing recovers 14 high-coherence biological features tracking KEGG pathways, NCBI taxa, and CARD AMR genes (r = 0.68–0.74); zeroing the butyrate-producer feature drops IBD AUC from 0.947 to 0.928. We conclude that multi-omic, multi-scale modeling is tractable at 7B parameters, and we release the model card, training-data manifest, evaluation pipelines, and a 100-trajectory MetaOmics-10T causal-pilot dataset.