[Mohamed Wahib] From Tokens to Tasks: Efficient Execution of Agentic AI at Scale
Abstract
Agentic AI shifts the performance challenge from serving individual model requests to executing dynamic workflows that interleave inference, tool execution, API calls, and data access. Optimizing inference alone is therefore insufficient to minimize task completion time. This talk presents our ongoing investigation into efficient execution of agentic workloads on heterogeneous computing platforms. We examine how CPU–GPU interactions, memory and I/O demands, and contention across concurrent tasks shape end-to-end performance, and explore scheduling and resource allocation strategies to improve throughput and latency. Using supercomputing platforms as a testbed, we investigate how these challenges evolve at scale and discuss their implications for the execution and resource management foundations of an agentic operating system.