Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering with Generative Optimization
Dapeng Jiang ⋅ Yizhe Chi ⋅ Kaisen Yang ⋅ Tianwei Luo ⋅ Boshi Zhang ⋅ Deyao Hong ⋅ Dianqiao Lei ⋅ Yifan Zhou ⋅ Weiyang Jin ⋅ Xiaoyan Fan ⋅ Han Hao ⋅ Zhe Cao ⋅ Qian Houde ⋅ Qingle Liu ⋅ Bohan Lyu ⋅ Bowen Wang ⋅ Situ Wang ⋅ Youjie Zheng ⋅ Bingxiang He ⋅ Eren Cai ⋅ Calvin Xiao ⋅ Qinhuai Na
Abstract
Engineering value is often created not by producing a single correct answer, but by improving a feasible artifact under hard constraints. We introduce **Frontier-Eng**, a benchmark for generative optimization agents that repeatedly edit executable engineering artifacts, receive frozen verifier feedback, and maximize the best feasible score found within a fixed budget. Frontier-Eng contains $47$ tasks across five engineering categories, spanning kernels, cryptography, quantum circuits, inventory and scheduling, robotics and control, optics and communication systems, structural design, reactions, and sustainable infrastructure. Each task provides a feasible starting solution, an editable artifact, and a read-only evaluator with verifier-parsed scoring, enabling cross-domain comparison without reducing engineering quality to binary pass/fail completion. We evaluate frontier models under OpenEvolve and compare search frameworks across the full task set. Current agents exhibit clear iterative improvement: **gpt-5.4** obtains the best average rank under OpenEvolve. Long-horizon runs reveal a dual **inverse-law pattern**: improvement events become rarer with iteration, and improvement magnitudes shrink with improvement count. At the same time, improvement is uneven across domains and remains sensitive to search depth, exposing a concrete frontier for engineering-oriented agents under realistic constraints and verifier-grounded feedback.
Chat is not available.
Successful Page Load