LLMs Optimizing LLMs: Automated MegaKernel Generation for Inference Acceleration
Weiqiang Xiong ⋅ Shaohui Peng ⋅ Wenyi Li ⋅ Hao Lu ⋅ Qirui Zhou ⋅ Congying Ma ⋅ Ziming Ye ⋅ Zhenyu Yi ⋅ Yunji Chen ⋅ Qi Guo ⋅ Ling Li
Abstract
High-performance inference for large models is critical in latency-sensitive applications. Existing inference engines suffer from sequential execution of hundreds of fine-grained kernels, causing kernel launch overheads and redundant memory accesses. A promising solution is fusing multiple kernels into a single large kernel, namely MegaKernel. However, existing MegaKernel implementations are tightly coupled to specific models and GPU architectures and require extensive manual engineering. We propose an LLM-driven framework AutoMegaKernel, which leverages an Instruction-Centric Fusion Abstraction to decompose MegaKernels into tractable fusion instructions, thereby enabling automated generation. Our framework jointly performs fusion strategy reasoning and CUDA code generation while hierarchically verifying correctness. Extensive experiments demonstrate that AutoMegaKernel consistently outperforms state-of-the-art inference engines (up to 5.04$\times$ vLLM and 2.42$\times$ SGLang) and compilers (up to 1.92$\times$ MPK), while supporting a broader range of models (dense, MoE, and VLA) and NVIDIA GPU platforms (Hopper, Ampere, and Ada).
Chat is not available.
Successful Page Load