SBIBM 2: A Simulation-Based Inference Benchmark Across Scientific Domains
Abstract
Inferring the parameters of complex stochastic simulators from empirical data is a key challenge across the sciences. Modern simulation-based inference (SBI) methods address this by reframing Bayesian inference as a supervised learning problem. Such SBI workflows are increasingly executed by agentic systems. Rigorously evaluating these approaches requires public benchmarks. An earlier SBI benchmark had standardized evaluations, but modern methods now increasingly saturate its low-dimensional tasks and risk pushing evaluation towards individual case studies that often lack a validated ground-truth posterior. Here, we present SBIBM 2, a suite of 24 SBI benchmarking tasks covering diverse scientific domains from epidemic dynamics to gravitational-wave astronomy, several of which demand high-fidelity inference to enable discovery in genuinely low signal-to-noise regimes. The benchmark is versioned, extensible, and openly hosted as a ready-to-use evaluation pipeline, with a backend-agnostic API, multiple data modalities, and a suite of evaluation metrics. Every task ships with a validated reference posterior and is assigned to a tier according to how that reference is obtained, from analytical through numerical to approximate. Baseline experiments reveal a clear performance gradient across tiers, pointing to substantial headroom for new methods. As evaluations depends only on the posterior estimates returned by a method, SBIBM 2 can be used both to assess individually configured algorithms and automated systems that assemble inference pipelines end-to-end.