FrontierOR: Benchmarking LLMs' Capacity for Efficient Algorithm Design in Large-Scale Optimization
Abstract
Large language models (LLMs) are increasingly used for optimization modeling and solver-code generation, yet practical operations research often requires a harder capability: designing scalable algorithms that exploit problem structure and outperform direct formulation-and-solve baselines. Existing benchmarks are limited to small or simplified examples far below real-world scale and complexity. We introduce FrontierOR, the first large-scale benchmark targeting LLM-based efficient algorithm design for realistic optimization problems. FrontierOR includes 110 tasks derived from methodologically diverse papers published in top-tier operations research venues, each with standardized instances and a hidden, human-verified evaluation suite. We evaluate seven frontier LLMs in one-shot and test-time evolution settings. The results reveal that frontier models still struggle to move from executable formulations to efficient optimization algorithms: the strongest one-shot model outperforms Gurobi in only 39\% of cases in terms of both solution quality and runtime, and even strong coding agents with test-time evolution achieve only 50\% on selected hard tasks. FrontierOR establishes a practical evaluation platform for LLM-based optimization algorithm design, which enables future LLMs and agents to be systematically tested on whether they can move beyond correct formulation toward feasible, high-quality, and runtime-competitive algorithms.