Grouped Adaptive Head Mixing for Personalized Multi-Task Federated Reinforcement Learning
Abstract
Multi-task reinforcement learning (MT-RL) trains a single agent to solve multiple tasks by leveraging shared knowledge across task domains. However, most MT-RL methods assume centralized access to task data, making them impractical for privacy-sensitive or large-scale distributed settings. Multi-task federated reinforcement learning (MT-FRL) mitigates this issue by enabling agents to collaborate through shared models rather than raw trajectories, but often suffers from unstable performance under task heterogeneity, negative transfer, and imbalanced learning across tasks. A fundamental question in MT-FRL is how each agent should decide with whom to share skills and how strongly to rely on shared knowledge. To answer this question, we propose GAdapHeadFRL, a personalized MT-FRL framework based on geometry-aware grouping and adaptive decision-layer task-head mixing. First, we decompose the local Bellman residual into representation and task-head estimation errors, and derive a closed-form optimal sharing strength governed by task mismatch and estimation uncertainty. Second, we instantiate this insight by identifying compatible client groups from local updated geometry, constructing group-specific task-head prototypes, and dynamically mixing them with local heads using lightweight practical proxies. Experiments on heterogeneous benchmarks, including MiniGrid (up to 12 tasks) and MetaWorld (up to 50 tasks), show that GAdapHeadFRL consistently outperforms state-of-the-art personalized MT-FRL baselines. It achieves up to 24.5\% relative improvement on MiniGrid, and increases the converged-task ratio from 50.0\% to 72.9\% over the personalized baseline. In addition, the proposed method reduces performance standard deviation by over 82.9\% on the largest task settings compared with centralized MT-RL references. These results demonstrate that geometry-aware grouping and adaptive task-head mixing provide an effective and scalable principle for heterogeneous MT-FRL.