PTXBench: Evaluating LLMs for GPU Kernel Optimization with Architecture-Specific PTX
Abstract
We introduce PTXBench, an evaluation framework for large language models' (LLMs') ability to use architecture-specific PTX to optimize GPU kernels. PTXBench measures functional correctness, whether selected target instructions execute at runtime, and speedup over frontier libraries on H100 and B200 GPUs. Our evaluation shows that architecture-specific PTX capability remains uneven across GEMM and attention workloads, and executing the target instructions does not necessarily translate into competitive performance. No evaluated recent LLM consistently matches frontier libraries across the suite. Controlled studies show that execution feedback improves the performance frontier on average. Under matched multi-turn refinement, Triton still generally reaches a higher best speedup and requires fewer reasoning tokens than CUDA--PTX but has exceptions. Under the same profiling budget, multi-turn refinement is more token-efficient than general-purpose coding agents, although Codex can overtake its matched multi-turn baseline on some workloads but not all after consuming more tokens. Together, these results position PTXBench as an extensible, auditable basis for measuring LLMs' ability to exploit evolving GPU architectures using PTX.