Stratagent: Modality-Specialized Agents Across the Genome, Transcriptome, and Proteome for Tumor Identity, Grade, and Survival
Abstract
Clinical characterization of a newly diagnosed tumor centers on three questions: what type of tumor it is, how aggressive its biology is, and what outcome the patient faces. We hypothesize that modality-specialized agents operating independently on genomic, transcriptomic, and proteomic evidence can characterize cancer more accurately than general-purpose language models reasoning directly over the same molecular inputs. Precisely, the Genomic Agent interprets copy-number alterations, driver mutations, tumor mutational burden, a microsatellite-instability proxy, aneuploidy, and whole-genome duplication; the Transcriptomic Agent evaluates gene-expression programs; and the Proteomic Agent analyzes total and phosphorylated protein abundance. An Orchestrator integrates these outputs using task-specific reliability weighting, detects disagreement with Jensen--Shannon divergence, and widens uncertainty or abstains when evidence conflicts. The language model performs tool selection, sequencing, reconciliation, and reporting, while quantitative predictions are produced by calibrated registry models. On 100 held-out TCGA cases, Stratagent achieved a 32-class tumor-identity macro F1 of 0.792 versus 0.223--0.329 for GPT-4o, GPT-5, and Sonnet 4.6. External evaluation on 100 CPTAC cases preserved the identity advantage, yielding a macro F1 of 0.792 and balanced accuracy of 0.750, whereas grade and survival did not separate the methods at that sample size. Fusing molecular layers did not exceed the strongest single layer, and reliability weights estimated on TCGA degraded under external platform shift until corrected by label-free reweighting. This paradigm reframes multi-omic cancer reasoning as coordinated specialization across molecular layers and suggests a path toward interpretable, uncertainty-aware tumor characterization.