Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning
Abstract
Real-world multi-agent reinforcement learning (MARL) systems must operate autonomously while adapting to natural language instructions provided by humans. These instructions arrive unpredictably, interrupt ongoing actions, and may conflict with long-horizon objectives under partial observability. Recent vision-language models enable instruction following, but are computationally expensive and limited to single-agent settings. Conversely, MARL can handle long-horizon coordination with macro-actions, but assumes uninterrupted execution. This work introduces macro-action value cancellation for instruction compliance (MAVIC) to enable interruption-aware MARL by bootstrapping the value of ongoing macro-actions upon receiving instructions. Agents can thus decouple task objective optimization from instruction-driven overrides within a unified policy and re-plan without destabilizing the underlying learning process. MAVIC integrates with standard actor-critic methods and achieves high instruction compliance while preserving base task performance across increasingly complex macro-action benchmarks.