Automating ML for Science: Can Frontier Agents Climb Scientific Hills in the Wild?
Abstract
Frontier coding agents are increasingly framed as end-to-end automated scientific researchers, with the promise of accelerating discovery on open-ended problems that lie beyond routine ML pipelines. We introduce Agent4Sci, a long-horizon testbed for measuring how close today's agents are to that vision, spanning 20 scientific benchmarks across 11 domains. Rather than ranking agents on a leaderboard, Agent4Sci asks what kinds of scientific work agents can automate under in-the-wild conditions, and where they still fall short. Empirically, today's frontier agents advance the published best on every benchmark, with Claude Opus reaching 10.4\% mean relative improvement and no run flagged for academic misconduct. These gains, however, come almost entirely from a bag of ML-engineering tricks rather than research-level insight. Trajectory analysis surfaces three recurring limitations in this setting: novelty miscalibration, resource under-utilization, and rapid collapse into local engineering sweeps. Lightweight plug-in skills for research ideation, experiment management, and scheduled checklists raise Claude Opus from 10.4\% to 20.1\% and Codex from 4.4\% to 8.1\%, while yielding more diverse, domain-relevant interventions. The gap to scientific research remains open. Overall, Agent4Sci gives the community a way to track and push frontier agents toward an automated researcher that proposes insights and drives measurable progress on real-world scientific problems.