AccelEval: A Whole-Program Benchmark for LLM-Generated CPU-to-GPU Code Acceleration
Chendong Song ⋅ Tianci Bu ⋅ Meixuan Wang ⋅ Zijie Zhou
Abstract
Current LLM-for-GPU benchmarks largely measure CUDA fluency on isolated operators or short kernel-generation tasks, whereas practical GPU performance engineering often starts from trusted sequential CPU programs and requires correct, end-to-end acceleration. To fill this gap, we introduce \textsc{AccelEval}, a benchmark for whole-program CPU-to-CUDA acceleration. AccelEval contains 42 tasks derived from public repositories and established HPC/application suites across six domains, each equipped with deterministic correctness oracles and three input scales. Beyond pass/fail correctness, AccelEval reports end-to-end speedup over CPU baselines and a novel peak-attainment metric which normalizes each model’s speedup against the best correct solution. Using this framework, we evaluate eight frontier LLMs across multiple input scales, GPU architectures and sample rates. At pass@1 on NVIDIA H200 at medium scale, the strongest model passes 38/42 tasks at 71.2$\times$ geometric-mean speedup over the CPU baseline. On the subset of 11 tasks with human-written CUDA references, the best-of-8 LLM oracle reaches 0.75$\times$ of human CUDA performance. To make CUDA-optimization behavior interpretable and reusable, we decompose correct solutions into 43 CUDA acceleration patterns and show that feeding winner-derived pattern guidance back to models improves paired geometric-mean speedup by 1.78$\times$, compared with 1.18$\times$ for a generic-tips control. Together, these results establish AccelEval as a reproducible diagnostic platform for end-to-end LLM-assisted GPU acceleration. An anonymized open-source release is available at \url{https://anonymous.4open.science/r/AccelEval-2397}.
Chat is not available.
Successful Page Load