Independent Data-Integrity Verification for LLM-Driven Functional Coverage Closure in Hardware Design
Abstract
Large language model (LLM) agents are increasingly used to close functional coverage gaps in hardware verification by generating directed test stimulus, but functional coverage alone only proves that a design state was reached, not that the stimulus which reached it produced correct data. We present a coverage-closure agent that treats these as separate concerns: an LLM iteratively generates directed cocotb tests against real RTL, with every coverage claim independently re-parsed from real Verilator simulation output rather than trusted from the model's self-report, while a separate reactive scoreboard monitor checks data integrity in real time against whatever stimulus is actually driven, rather than only a disconnected fixed sequence. Across three architecturally distinct RTL designs (an asynchronous FIFO, a UART transmitter, and an AHB2APB bridge), the agent closes functional coverage from a random-stimulus baseline to 100% using unmodified, design-agnostic logic, and the reactive monitor independently confirms zero data-integrity failures in the base stimulus while uncovering a genuine, previously undetected clock-domain-crossing bug in the testbench's own stimulus logic. We report this result together with an explicit statement of what remains unverified: the correctness of the LLM's own generated directed-test stimulus is not yet independently checked in this version, motivating ongoing work on a bounded, safeguarded verification pipeline for the generated stimulus itself.