HDL-RepoBench: Multi-Paradigm Repository-Level Code Completion for Hardware Design Languages
Abstract
Large language models (LLMs) have achieved strong performance on code completion in general-purpose programming languages, but existing repository-level benchmarks focus almost exclusively on software code and largely overlook hardware description languages. We present HDL-RepoBench, consisting of HDL-RepoBench-Train and HDL-RepoBench-Eval, the first benchmark for multilingual hardware code completion at the repository level with functional evaluation. HDL-RepoBench covers three major hardware design coding styles and annotates each completion target with code-structure-level and hardware-oriented semantic labels derived from concrete syntax tree analysis. Post-training billion-scale code LLMs on HDL-RepoBench-Train enables smaller open-weight models to outperform much larger general-purpose models under matched context windows. Beyond aggregate accuracy, our analysis surfaces four findings that distinguish hardware from software code completion: (i) performance varies sharply across HDLs, with HLS consistently easiest and VHDL hardest across functional pass rate and EM/ES; (ii) hardware accuracy is non-monotonic in code-structure depth, in contrast with the monotonic depth–accuracy curve known for software; (iii) accuracy plateaus once the input context exceeds 2,048 tokens, while software completion continues to benefit from longer context; and (iv) certain repository-level retrievers tuned on software underperform a no-retrieval baseline on HDL repositories because their structural and dependency assumptions are violated by hardware code.