Drift-Aware Meta-Agent Training for Long-Horizon LLM Multi-Agent Coordination
Abstract
Large language model (LLM)-based multi-agent systems have shown strong collaborative capabilities through communication, planning, and role specialization. However, improving individual decisions does not necessarily ensure stable team behavior over long-horizon interactions. Under partial observability, locally plausible decisions can become mutually incompatible over time, resulting in conflicting sub-goals, redundant work, and progress-free execution loops. We refer to these recurring failures as coordination drift. We propose DriCo, which uses coordination drift to learn and revise a shared team-level context. DriCo trains a coordinator to prefer contexts that induce compatible and non-redundant sub-goals. During execution, it selectively updates the context and affected agents' sub-goals when drift emerges. The resulting context guides LLM-based planners and Q-guided actors for hierarchical agent-level execution. We further introduce LLM-Overcooked with diverse layouts and recipes for evaluating long-horizon coordination. Experiments show that DriCo improves performance while reducing unresolved conflicts, redundant sub-goals, and execution loops. These results demonstrate the benefit of learning and revising shared coordination context from observable coordination failures.