HPC-Bench: A Comprehensive Benchmark for High Performance Computing Codes
Bowen Cui ⋅ Junyu Yin ⋅ Tejas Ramesh ⋅ Leo Lim ⋅ Oscar Hernandez ⋅ Keren Zhou
Abstract
Large Language Models (LLMs) have shown strong capabilities in code generation and reasoning, motivating extensive benchmarks for evaluating functional correctness. However, performance optimization in High-Performance Computing (HPC) poses fundamentally different challenges that are not captured by existing evaluations: HPC performance depends on low-level code, intricate memory access patterns, and computational kernels embedded within sophisticated execution contexts. We introduce HPC-Bench, a performance-grounded benchmark for evaluating LLMs as optimizers of computational kernels in realistic HPC workloads, including $114$ workloads across $16$ HPC computational motifs. HPC-Bench provides a fully automated, end-to-end evaluation pipeline that combines program-level correctness verification with two correctness-gated metrics, $\mathrm{Fast}@k$ and $\mathrm{Speedup}@k$, to characterize reliability and speedup under a fixed sampling budget. Our evaluation of three representative LLMs across serial CPU, OpenMP, and CUDA settings shows that stricter speedup requirements substantially reduce the probability of finding a correct and fast optimization. HPC-Bench is available at https://github.com/neuripspaper2026/NeurIPS_26.
Chat is not available.
Successful Page Load