PCBInnoBench: Benchmarking LLM Agents on Real-World PCB Design
Abstract
LLM-based agents have made rapid progress on software engineering benchmarks, and executable evaluation is extending to chip-level hardware tasks. Board-level PCB design and upgrade has yet to receive comparable evaluation, largely because open-source PCB projects rarely provide the structured issue/PR/test histories that existing benchmarks rely on. We introduce PCBInnoBench, the first benchmark that evaluates LLM agents on executable PCB design-and-upgrade workflows. PCBInnoBench contains 372 expert-designed KiCad tasks grounded in real PCB projects and end-to-end engineering workflows. Each task pairs an agent-visible package (KiCad project and multimodal engineering references) with a hidden evaluation package (ground-truth design and expert-reviewed validators) used only for scoring. Agents must modify existing ECAD projects and pass structural, task-specific, and cross-artifact engineering-constraint checks. Each task is evaluated under two prompts of different guidance levels: one supplies an expert execution trajectory of ECAD edits together with the consulted references, the other states only the task objective, requiring agents to navigate the full workflow from requirement localization through cross-artifact editing. The best agent achieves 22.0% Pass@1 with guidance but only 7.3% without, indicating that both workflow discovery and ECAD execution remain substantial challenges for current agents. Beyond PCB, the workflow-grounded, expert-constructed evaluation approach may serve as a reference for other engineering domains where historical traces are sparse and correctness spans heterogeneous artifacts. Data and code are available at https://anonymous.4open.science/r/pcbinnobench-C841 .