AI-Ready Metagenomics: Representation Choice Shapes Access to Clinical Treatment-Response Signal
Abstract
Scientific AI can fail before model training begins: transforming raw experimental measurements into analysis-ready representations may discard information required for the downstream task. Shotgun metagenomics provides a particularly clear example because raw sequencing measurements are routinely reduced to known taxa and pathways despite substantial incompletely characterized sequence space. We used treatment response in Crohn's disease as a real-world case study of this representation problem, analyzing baseline stool metagenomes from 89 patients initiating anti-TNF, vedolizumab, risankizumab, or ustekinumab treatment. The same quality-controlled reads were processed through taxonomic, metabolic-pathway, and reference-independent 40-mer co-occurrence representations. The reference-independent procedure generated treatment-specific reference-independent scores (RIS), with dictionary construction, biomarker discovery, normalization, feature selection, and score definition performed within training partitions only before application to held-out patients. RIS discriminated composite remission across all four therapies (AUC 0.801-0.955), whereas taxonomic and metabolic-pathway models derived from the same metagenomes showed lower and inconsistent direction-corrected AUCs and no significant global remitter/non-remitter separation by PERMANOVA. Because the downstream procedures are not identical, these comparisons are descriptive rather than a formal algorithm benchmark. The findings illustrate that scientific data may be clean, curated, and technically analysis-ready while still being information-poor for AI because of choices made at the representation layer. They also highlight a second data-readiness requirement: when a representation is learned from the cohort, all adaptive representation construction must remain inside the training boundary before held-out evaluation.