Beyond Mutation Coverage: Auditing Verdict Reliability in RTL Generation Benchmarks
Abstract
Most RTL generation benchmarks assess functional correctness from testbench verdicts. Prior work has audited the mutation adequacy of these testbenches and, separately, reported that a substantial fraction of testbench-accepted RTL fails formal equivalence checking. However, the relationship between official testbench adequacy and actual verdict errors on LLM-generated RTL --- including both erroneous candidates that are accepted and correct candidates that are rejected --- remains underexplored. We reproduce the published mutation audit of the RTLLM v2.0 benchmark and construct an independent reference-differential oracle to evaluate 690 outputs from two open-weight LLMs. Incomplete mutation detection is associated with functionally incorrect designs passing the official testbench (14.5\% of accepted candidates), whereas positional interface assumptions reject behaviorally correct candidates even for tasks with a perfect mutation score. We identify vacuous execution and positional binding as two distinct failure modes that mutation adequacy alone cannot capture, and release the evaluation artifacts.