VeriScope: Measuring Verification-Ready Verilog Artifacts
Wei Zhang ⋅ Jian Yang ⋅ jiajun wu ⋅ Junhang Cheng ⋅ Chufan He ⋅ Yihang Lou ⋅ Xianglong Liu
Abstract
Large language models (LLMs) are increasingly used in hardware-oriented coding workflows, yet open resources for evaluating first-pass RTL beyond execution remain limited. We present \benchmark, an open benchmark and evaluation suite of \textbf{568 problems} spanning basic combinational gates through module-level designs and a 95-task L4+ top-end slice (92 L4 system tasks plus 3 L5 stress tests). A developer agent receives a natural-language task brief, returns a single RTL draft, and that artifact is evaluated through objective execution plus RTL-level and artifact-level review incorporating waveform and circuit evidence. Across \textbf{26 public models} plus two released 32B calibration baselines, the best systems reach about 84/100 combined score on the full benchmark but drop to the high 70s on L3--L4+. The 60/40 combined leaderboard largely tracks functional execution (Spearman 0.998 with a functional-only ranking), while artifact scores are most useful as a post-simulation screening signal: 29.6\% of L3--L4+ simulation-passing submissions receive artifact score below 6 and therefore need additional verification evidence. A 100-sample manual taxonomy of such disagreements finds \textbf{70\%} testbench coverage gaps, \textbf{12\%} real RTL defects, and \textbf{18\%} judge hallucinations (Wilson 95\% CIs roughly $\pm 9$\,pp), showing that artifact review surfaces both under-tested passes and model defects, but should not be treated as a standalone leaderboard refinement.
Chat is not available.
Successful Page Load