Beyond Equivalence: Task-Level Verification Unlocks Architectural Optimizations by LLM Agents
Abstract
Automated RTL optimization is conventionally verified by logical equivalence. The optimized design must reproduce the reference's cycle-by-cycle I/O behavior, which fixes the schedule and the memory access pattern. We relax the verification contract to the task level, where the only requirement is correctness of the output data for given input data, leaving the cycle schedule, memory access order, and internal architecture free. On six benchmarks ranging from matrix multiplication to a RISC-V ALU, an LLM optimization agent improves the baseline RTL design under both formulations. As objectives we use cycle-aware cost metrics on a fixed workload: the area–runtime product and its energy-weighted extension. Equivalence-constrained optimization reduces the area–runtime product by up to 62%, while task-level freedom reaches up to 82%. However, the winning design may consume more energy than the baseline. On matrix multiplication, where this occurs, we show that energy-directed optimization yields leaner designs that use 28% less energy than the baseline at competitive area–runtime product. A further analysis on the matrix multiplication benchmark weights the cost by data-movement count and shows that an exponent on the movement term acts as a dial over the accelerator design space, steering the optimal architecture from a single multiplier through operand-reuse tiles to a fully buffered parallel array.