SiliconEye: A Multi-modal Understanding Benchmark for Hardware Design
Sriram Ranga ⋅ Mehmet A Korkmaz ⋅ Aneesh Bendale ⋅ Pranav Lakshminarayan ⋅ Anupam Chattopadhyay ⋅ Debjyoti Bhattacharjee
Abstract
Today's multi-modal models are increasingly used for assistance on a variety of hardware digital design tasks, yet they are measured almost exclusively on text-only benchmarks. Modern hardware implementation is difficult to interpret from code alone because its behavior spans modules and clock cycles. Block diagrams, waveforms, and state-transition graphs reveal complementary structural, temporal, and control-flow information. Current multi-modal hardware benchmarks inadequately measure models’ ability to interpret figures from complex real-world design artifacts, particularly research papers. In this work, we address this gap by introducing SiliconEye, a multi-modal understanding benchmark. SiliconEye offers a training dataset sourced from over around 80k Verilog-primary GitHub repositories, consisting of more than 6000 relevant visual hardware design related images, tagged with relevant metadata, specification documents and research papers. Each sample in the benchmark is triangulated by the design image, paper text and the reference RTL. From these, we manually curate $84$ objectively gradable MCQ, numerical and yes\no type questions spanning block diagrams, architecture images, FSMs, waveforms among others. We compare various models under different conditions and observe that the accuracy ranges from $58.3\%$ for a small open source model with thinking disabled to $86.9\%$ for a frontier thinking-enabled model. To evaluate whether models rely on visual evidence and whether they can answer correctly when that evidence is barely readable, we additionally conduct experiments using degraded images and a judge model.
Chat is not available.
Successful Page Load