AgentWeave: Efficient Distributed Agent Serving via Flow Decomposition
Kaibin Guo ⋅ Pengtu Li ⋅ Zicong Hong ⋅ Zhiyuan Fang ⋅ Wuhui Chen ⋅ Rachid Guerraoui ⋅ Anne-marie Kermarrec
Abstract
Agent applications are evolving into workflows composed of large language model (LLM) calls and tool invocations, termed modules. Existing distributed agent systems typically adopt a disaggregated architecture, deploying different modules on separate devices. While this enables pipeline parallelism, imbalanced module latencies introduce severe pipeline bubbles that limit throughput. A natural alternative is an aggregated architecture that co-deploys all modules across all devices, eliminating bubbles by keeping every device continuously active. However, this introduces request blocking: short requests are forced to progress in lockstep with longer ones in the same batch, and newly arriving requests must wait until the entire ongoing batch completes the full workflow. To address this, we propose AgentWeave, an efficient distributed serving system for agent applications built on the aggregated architecture. Its core mechanism, flow decomposition, decomposes long agent requests into fine-grained execution units while preserving agent semantics, allowing short and newly arriving requests to proceed without being stalled by long-running ones. Complementing this, a flow-adaptive scheduler dynamically balances the throughput gains of decomposition against its scheduling overhead. Experiments across four representative agent applications demonstrate that AgentWeave improves throughput by $1.81\times$, reduces average latency by $2.45\times$, and reduces P90 latency by $2.04\times$ over state-of-the-art baselines.
Chat is not available.
Successful Page Load