CHoRD: Coordinating Scheduling and Data Placement for Efficient Deep Neural Network Inference on Chiplet-Based GPUs
Abstract
Chiplet-based GPUs are a promising alternative to monolithic designs for scaling compute and memory resources in machine learning (ML) workloads. However, distributing computation and memory across multiple chiplets introduces a hierarchical Non-Uniform Memory Access (NUMA) environment, where performance depends on the alignment between kernel execution and data placement. In deep neural network (DNN) inference, kernels execute repeatedly over shared weights and activations, and poor alignment leads to sustained inter-chiplet traffic that degrades performance. Existing approaches optimize scheduling or data placement in isolation, failing to coordinate these decisions across dependent kernel executions. We present CHoRD (Chiplet-Hierarchy-oriented Residency and Dispatch), a compiler-assisted, profile-guided framework for joint scheduling and data placement in chiplet-based GPUs. CHoRD reconstructs program structure from compiled kernels and host code, and supports two modes: a static mode that derives scheduling and placement decisions without profiling, and a profiling mode that performs a single lightweight run to collect execution statistics. Using this information, CHoRD determines chiplet-level kernel mapping, enables safe inter-kernel overlap, improves intra-chiplet parallelism, and assigns data to chiplets based on access patterns. These coordinated decisions reduce inter-chiplet communication and improve compute utilization. We evaluate CHoRD on 4-chiplet and 8-chiplet MGPUSim substrates across nine CNN and transformer-based inference workloads. On the 4-chiplet system, CHoRD-static and CHoRD-profiled achieve 2.57× and 2.94× geometric-mean end-to-end speedups over RR, and 1.46× and 1.66× over the placement-only CLAP baseline, respectively, with benefits increasing as chiplet count scales to 8.