Efficient Agentic GPU Kernel Optimization with a Compact Domain-Specific Language and Speed-of-Light Guidance
Siva Kumar Sastry Hari ⋅ Vignesh Balaji ⋅ Sana Damani ⋅ Qijing Huang ⋅ Christos Kozyrakis
Abstract
LLM agents can optimize GPU kernels, but each candidate requires generation, compilation, correctness testing, and profiling. Our objective is to reduce the number of expensive attempts needed to find a correct, fast kernel. We study two complementary ways to improve this attempt efficiency by changing the representation that the agent writes in and adding a first-principles performance signal that estimates remaining headroom. We introduce $\mu$CUTLASS, a compact in-context learnable DSL that exposes high-impact CUTLASS optimization choices while hiding template plumbing, and pair it with Speed-of-Light (SOL) guidance for search steering, budget allocation, and integrity checking. On 59 KernelBench problems on H100, switching GPT-5-mini from low-level code generation to $\mu$CUTLASS changes a $0.40\times$ geomean regression versus PyTorch into a $1.27\times$ speedup. Adding SOL-guided steering, it reaches $1.56\times$. Across model tiers, this combination lets weaker models match or exceed stronger raw-code baselines, and SOL-guided scheduling saves 19--43\% of tokens while retaining at least 95\% geomean speedup. We also show that integrity filtering is essential because, without SOL-assisted checks, reported speedups can be inflated by up to $1.9\times$ by benchmark-gaming or PyTorch-only solutions.
Chat is not available.
Successful Page Load