ML4SCIENCE-BENCH: Benchmarking Autonomous Agents on End-to-End Machine Learning for Science
Abstract
Coding agents can write and execute code, but completing a scientific machine-learning task requires more than producing an executable program. An agent must construct a valid end-to-end workflow: prepare data correctly, prevent evaluation leakage, select appropriate models and evaluation procedures, respect domain-specific scientific constraints, and produce evidence that supports its conclusions. We introduce ML4SCIENCE-BENCH, a benchmark containing 100 executable tasks inspired by scientific papers across biology, chemistry, mathematics, and physics. Each task requires an agent to develop, analyze, or examine a complete machine-learning workflow in a terminal environment. Task-specific deterministic tests not only evaluate execution outcome, but also ML logic, scientific consistency, and reproducibility. We evaluate 10 language models, using a common TERMINUS-2 scaffold with all models configured for high reasoning effort. Despite access to the same environment and tools, The fraction of runs passing every test ranges from 24.7\% to 77.3\%. Average cost per run varies by more than sixfold, and higher success rates do not consistently correspond to higher cost. These results reveal a substantial gap between generating executable code and carrying out scientifically valid machine-learning workflows, and show that agent evaluation in scientific ML must jointly consider correctness, reliability, cost, and execution time.