ASE-Bench: An Atomic Simulation Benchmark for Agentic Materials Discovery
Abstract
Large language models (LLMs) have become increasingly capable of writing the code that drives atomistic simulation. Benchmarks track their progress in software engineering and mathematics, yet no benchmark measures whether an LLM can construct and execute an atomistic simulation correctly. ASE-Bench poses 100 curated Atomic Simulation Environment (ASE) tasks in five difficulty levels, spanning structure building, optimisation, molecular dynamics, vibrations, and thermochemistry, to 94 LLMs including frontier models from ChatGPT to Qwen. To evaluate whether in-context knowledge of ASE matters, each model is evaluated without and with a reference ASE skill file. Every generated script is executed and graded against quantitative simulation results. On the 50 easy tasks (levels L1-L3), the best model, gpt-5.6-sol, solves every task without the skill file. The ASE skill raises accuracy by 13.3% on average and by up to 43% for weaker models, and per-task cost varies several thousand-fold across models. On the 50 hard tasks (levels L4-L5), which include oxide surface reconstruction and heterojunction construction, no evaluated model exceeds 50%. This work establishes a standard for evaluating LLM-written atomistic simulation code. Current models handle most easy tasks, the hard tasks remain unsolved, and the full benchmark is released to track that progress.