Forecasting Microbial Dynamics: Evaluation Protocol and Prior-Spectrum Benchmark
Abstract
Modern forecasting methods are often designed and evaluated on datasets that do not reflect the realities of real-world clinical time-series: inter-patient variability, feature constraints, irregularity, and sparse sampling. Microbiome dynamics modeling represents these challenges, has seen little attention from the machine learning community, and has practical significance in medicine. Moreover, it offers both deep domain knowledge and sufficient data to train deep models, making it a strong testbed for a wide variety of forecasting methods. We propose a unified evaluation protocol for forecasting microbial dynamics and conduct MicroCastBench: a benchmark for models covering a wide range of priors (from mechanistic to foundation models) and assumptions (from per-patient to meta-learning models). We find that meta-learning models lead across most scaling regimes and at longer horizons, with mechanistic priors winning at small cohort sizes, short trajectories, and high feature counts. These results suggest that when the target data is far from pretraining distributions, domain-informed and meta-learning approaches can outperform generalist foundation models, highlighting the value of evaluating on under-represented domains.