Know Your Task, Learn It Right: Task-Aware Optimistic Value Learning for Multi-Task Multi-Agent Reinforcement Learning
Abstract
Cooperative multi-agent reinforcement learning under task variability requires both identifying which task an episode belongs to and adapting the learning algorithm to that identity. Recent class-aware methods address the first half of this problem by clustering trajectories into latent task classes and feeding the class label to the agent policy. They leave the second half largely untouched: the Bellman update and the credit assignment performed by the value mixer remain identical across all task classes. We show empirically that, even when task identity is recovered with high accuracy, the Bellman dynamics across classes diverge sharply, with simple classes saturating early and harder classes sustaining large temporal-difference residuals and large cross-agent value disagreement. Building on this observation, we propose Know Your Task, Learn It Right (KYT), a task-aware extension of value factorization that injects per-class learning signal into the Bellman update itself. KYT maintains a lightweight tracker that summarizes the per-class TD-error and cross-agent value spread without any extra network, and uses the tracker in two places: an optimistic target that adds a bounded per-class exploration bonus, and a class-conditioned mixer that modulates credit assignment through a softplus-gated affine layer preserving the IGM property. Across StarCraft II micromanagement benchmarks including SMACv2 and the unit-combination suite SurComb, KYT outperforms strong class-aware and exploration-aware baselines, with the largest gains concentrated on the hardest task classes that current methods underfit.