ASTOR: Multi-Task Code Reinforcement Learning via Utility-Driven Coordination
Abstract
Reinforcement learning (RL) with verifiable rewards has proven effective at post-training LLMs for coding, yet deploying separate task-specific specialists incurs costs that scale with the number of tasks, motivating a unified multi-task RL (MTRL) approach. However, existing MTRL methods treat all coding tasks uniformly, relying on fixed data curricula under a shared optimization strategy, ultimately limiting the effectiveness of multi-task training. To address these limitations, we propose \textbf{{\tool}}, a multi-t\textbf{AS}k code reinforcement learning framework via u\textbf{T}ility-driven co\textbf{OR}dination. Centered on \textit{task utility}, a signal capturing each task's learning potential and cross-task synergy, {\tool} comprises two coupled modules: \textit{1) \moduleone} module hierarchically allocates training budget and prioritizes informative prompts, steering training toward the most valuable data; and \textit{2) \moduletwo} module dynamically scales per-task KL regularization, matching update constraints to each task's current training state. Experiments on two widely-used LLMs across four representative coding tasks demonstrate that {\tool} consistently improves a single model across all tasks, outperforming the best task-specific specialist by 9.0\%--9.5\% and surpassing the strongest MTRL baseline by 7.5\%--12.8\%.