Resource-Aware Joint Kernel Fusion for Ahead-of-Time Inference
Abstract
Real-time inference of learned models in environments with low latency and memory consumption constraints is becoming challenging with their growing complexity. Hardware architectures such as GPUs are actively utilized for developing efficient serving systems, leveraging their massive parallelism benefits towards computation heavy operations for throughput improvements. Kernel fusion, being a method to harness parallelism in GPUs, is typically applied along a single axis, either horizontal (co-launching independent operators to maximize occupancy) or vertical (chaining dependent operators to avoid DRAM round-trips), with each axis unaware of the other's decisions, and sized against the hardware's physical ceiling rather than what the application is willing to spend. We present Jores, a joint fusion scheduler that builds vertical fusion candidates and pairs them with horizontal grouping, considering mutual benefits under a user-declared resource budget. We evaluate across four transformer-style building blocks, chosen such that every block contains both a horizontal and a vertical fusion opportunity. We compare against internal single-axis ablations and against policies approximating related work, reporting DRAM bytes saved, kernel launch count, latency, concurrency achieved, and scheduling overhead. We discuss extensions toward learned scheduling policies as future work.