MDArena: Evaluating Coding Agents on Realistic Molecular Dynamics Workflows
Abstract
Coding agents increasingly automate scientific workflows, but their reliability in molecular dynamics (MD) research remains unclear. To address this issue, we introduce MDArena, a benchmark of 50 containerized tasks drawn from active biomolecular simulation projects, spanning 29 molecular systems and 14 broad research protocols. We evaluate seven model-harness configurations spanning Codex, Claude Code, and OpenCode, with three independent attempts per task and 1,050 trials in total. Codex-GPT-5.6 Sol performs best, achieving an 82.0% Strict-Pass@1 rate (123/150 trials), followed by Claude Code-Opus-5 at 76.7% and OpenCode-Gemini 3.7 Flash at 71.3%. Correctness and process rewards consistently exceed strict success rates, indicating that agents frequently make meaningful partial progress but fail on the fine-grained details required for reproducible scientific workflows. These failures were especially prevalent on hard, long-horizon workflows. Across all hard tasks, the best-performing configuration was Claude Code-Opus-5, with only 52.1% strict success (25/48 trials), closely followed by Codex-GPT-5.6 Sol at 50.0\% (24/48). MDArena therefore highlights the current limits of coding agents and provides a reproducible and extensible platform for tracking progress toward closing this gap.