BioTrace: Disentangling literature bias from biological signal in discovery tasks
Abstract
Large Language Model (LLM)-based AI agents are increasingly evaluated for biological discovery, but strong performance may reflect recovery of well-established knowledge rather than identification of underexplored biology most relevant to the research questions being considered. Here, we developed BioTrace, a biological interpretability framework that characterizes model outputs across three axes of evidence: gene representation in the scientific literature (PubMed), connectivity in curated biological databases (Reactome, STRING), and connectivity measured through large-scale experimental screens (DepMap, HI-union, Hu.MAP). We applied BioTrace to five published studies (AssayBench, PerturbQA, GenePrior, GeneTuring, LLM-SynthLet) comprising 73 model and method configurations across diverse biological tasks related to gene prioritization. Publication history and curated connectivity were generally more strongly associated with model outcomes than screen-derived connectivity, a pattern also observed in non-LLM baselines. Model adaptation and scale altered these associations, and size-matched models from different families showed similar profiles, but none of these consistently shifted prioritization toward experimentally-derived connectivity. Adjusting for publication history attenuated associations with curated connectivity, particularly documented physical interactions. Among understudied genes, greater screen-derived connectivity could coincide with poorer recovery of experimentally supported hits. BioTrace thus provides a reusable framework for distinguishing recovery of established knowledge from prioritization of underexplored biological relationships.