DFT-Agent Bench: What Agent Engineering Buys, and What It Does Not, for Autonomous DFT
Shijie Liu ⋅ Jin Yang ⋅ Weilin Wan ⋅ Jingtao Han ⋅ Jia LI ⋅ Tianqi Deng
Abstract
Large language models are increasingly deployed as autonomous agents for first-principles materials simulation. Each new system wraps the model in more scaffolding: structured tools, knowledge layers, verifier loops, sub-agents. How much this engineering matters is rarely measured. We introduce \textbf{DFT-Agent Bench}, a benchmark of 200+ Quantum ESPRESSO tasks spanning relaxation, electronic structure, equation-of-state, and elastic-tensor pipelines over multiple material tiers. Each task carries a cross-validated ground truth, and a provenance audit voids any score not backed by a real, isolated calculation. Comparing a near-bare ReAct baseline against a heavily-engineered flagship on two frontier models, we find that the returns on engineering are conditional on baseline headroom. Where the baseline is already accurate (relaxation, equation of state), engineering adds nothing, and the gate even costs the flagship points for output-contract misses. Where the baseline has headroom, the gains are real and transfer across models: bands+DOS improves by $+0.08$/$+0.11$ ($p\!\le\!0.006$) and $d$-electron vc-relax by $+0.17$ on both models ($p\!\le\!0.009$). The elastic tensor improves only on the model whose baseline collapses ($0.18\!\to\!0.40$ on \texttt{deepseek-v4-pro}; the \texttt{claude-opus-4-8} baseline already reaches $0.62$). A frozen 100-task re-run with retrieval disabled (the deepseek baseline re-run frozen alongside) confirms the gains on both models ($+0.10$/$+0.17$ paired, $p{<}0.001$, parity $0.88$/$0.88$). An A/B shows no detectable contribution from the experience store: the gains come from curated scaffolding, not cross-task memory. The audit voids only $1$--$2\%$ of runs for both agents. The open frontier is task difficulty: no agent--model pair exceeds $0.62$ on the elastic tensor. We release the tasks, ground-truth pipeline, scorer, and both solvers (anonymized: \url{https://anonymous.4open.science/r/dft-agent-bench}).
Chat is not available.
Successful Page Load