Accelerating the Inference Era with AI-Driven, Globally Optimized HW/SW Co-Design
Miria Feng ⋅ Fangzhao Zhang ⋅ Adrian G Lafuente ⋅ Mert Pilanci ⋅ Azalia Mirhoseini
Abstract
Deep learning inference demand is projected to grow $10{,}000\times$ over the next five years, a trajectory that general-purpose accelerators cannot match. Custom accelerators offer a viable path forward, but designing them requires simultaneous co-optimization of hardware datapaths, software schedules, and compiler transforms across a combinatorial search space exceeding $\mathcal O(10^{2300})$. While co-design is becoming increasingly essential, existing methodologies rely on decoupled, sequential pipelines that miss critical cross-stage trade-offs. We introduce CHEETAH: a globally optimized, AI-driven framework for inference accelerator co-design. To our knowledge, this is the first co-design framework to combine \textit{joint} tiling-fusion, dynamic placement, and rematerialization in a single LP-based formulation. Key contributions include: (1) time-indexed variables for dynamic memory management, (2) rematerialization variables to enable cheap recomputation over costly DRAM reloads, and (3) joint tiling-fusion optimization via McCormick linearization. This inner-loop precision is steered by an LLM-guided outer loop that performs causal interventions based on specific deployment profiles and Lagrangian sensitivity signals. In less than 6 hours, \textsc{CHEETAH} discovers designs on the Pareto frontier achieving an average of $\sim22.07\times$ raw speedup across $11$ workloads, reaching up to $\sim40.1\times$ speedup on a single workload over a TPU-v3 baseline, unlocking efficiency inaccessible to decoupled sequential pipelines.
Chat is not available.
Successful Page Load