Evolving Agent Teams
Shiyi Cao ⋅ Ziming Mao ⋅ Dacheng Li ⋅ Joseph Gonzalez ⋅ Ion Stoica
Abstract
The rapid progress of software engineering agents drives interest in applying them to open-ended problems such as GPU kernel optimization. However, we find that existing coding agents converge to local optima on real-world, expert-written kernels (e.g., FlashInfer). We introduce Evolving Agent Teams, a multi-agent framework that escapes such local optima. It features (1) a central supervisor that diagnoses each local optima and restructures the roles of a team of agents; and (2) a persistent file system that comprises a strategy tree, an evolution log, and per-team artifacts, grounding each restructuring in accumulated experience. Under a matched token budget, Evolving Agent Teams consistently outperforms strong coding-agent baselines. With an extended token budget, Evolving Agent Teams generates a CUDA kernel that reaches $1.2\times$ mean speedup over the production FlashInfer MLA paged decode kernel and wins on 36 of 47 production workloads from FlashInfer-Bench ($0.86$--$0.99\times$ on the remaining 11).
Chat is not available.
Successful Page Load