DERBench: A Benchmark for Evaluating LLMs on Distributed Energy Resource Operations
Shreya Prithviraj Savant ⋅ Chengming Hu ⋅ Juanwei Chen ⋅ Kajal Pokharel ⋅ Yefeng Yuan ⋅ Jun Yan ⋅ Yuhong Liu ⋅ Hepeng Li ⋅ Jie Gao
Abstract
Distributed energy resources (DERs) turn distribution-grid operation into a coupled planning, monitoring, and control problem: an interconnection recommendation, a state estimate, or a dispatch action must be consistent with feeder conditions and operating limits. Large language models (LLMs) are increasingly being explored for these workflows, yet existing evaluations rely heavily on question answering and thus conflate knowledge with the ability to execute operational tasks. We introduce DERBench, a unified multi-tier benchmark for evaluating LLMs on DER operations across $174$ tasks, $9$ operational categories, and $4$ tiers: knowledge retrieval, contextual reasoning, executable modeling, and multi-turn agentic operation. Each task is evaluated using a scoring mechanism fixed before evaluation and matched to the task: answer keys for closed-form responses, sandboxed execution against solver-based power-flow and optimization references for executable tasks, environment-based scoring for agentic episodes, and dual-judge evaluation restricted to open-text criteria. We evaluate eleven frontier models from nine providers zero-shot. All models exceed $83.0%$ on knowledge retrieval, while L3 executable modeling spans $58.9–89.8%$ and reorders the field relative to retrieval. Agentic episodes achieve higher overall scores ($80.2–93.5%$), while uncertainty handling remains the weakest process dimension. We also find that evaluation design materially affects measured performance: human review of 482 disagreements across $4,797$ dual-judged criteria shows that unanimous judge agreement systematically under-credits open-text responses, shifting retrieval-tier means by $4.0$ points; repeated runs, prompt perturbations, and alternative grading rules leave model orderings largely stable. These results motivate tier-resolved evaluation for assessing LLMs before operational use. Dataset and code are available at https://anonymous.4open.science/r/DERBench.
Chat is not available.
Successful Page Load