Does Pretraining Pay Off?: A Systematic Comparison of Microbiome Foundation Models and Simple Baselines
Abstract
Foundation models pretrained on large microbiome datasets have been gaining interest as a way to transfer knowledge learned from structured microbial communities and facilitate tasks such as phenotype prediction. Numerous methods have been proposed with varied deep learning and language modelling inspired architectures, most notably scGPT-style masked transformer, MGM causal autoregressive transformer, and masked autoencoders. Contrary to reports on individual datasets, we find no consistent benefit from pretraining or transformer architectures over simpler tree-based models operating on raw or normalized taxa counts. We investigate whether this gap can be explained by data sparsity, low effective sample size after accounting for study-level batch effects, or predictive signal that is concentrated in a small number of taxa or even fully explained by technical confounds such as sequencing depth. Our results suggest that current microbiome foundation model approaches are not yet delivering the benefits seen in adjacent domains, and that data quality and evaluation rigor, not model capacity, are the limiting constraints.