Drift-Aware Coordination Learning for LLM-Based Multi-Agent Systems
Abstract
Large language model (LLM)-based multi-agent systems have demonstrated strong collaborative capabilities through communication, planning, and role specialization. Yet stable team behavior remains difficult over long-horizon interactions. Under partial observability, individually plausible decisions can gradually diverge, leading to conflicting sub-goals, redundant work, and progress-free execution loops. We term these recurring failures coordination drift. We propose DriCo, which learns a shared team-level context that promotes compatible, non-redundant sub-goals and selectively revises the context and affected agents’ sub-goals when drift is detected. The refined context guides hierarchical execution through LLM-based planners and Q-guided actors. We further introduce LLM-Overcooked for evaluating long-horizon coordination across diverse layouts and recipes. Experiments demonstrate that DriCo improves task performance while reducing conflicts, redundant sub-goals, and execution loops.