Balanced Multi-Task Learning from an Optimality-Gap Perspective
Abstract
Multi-task learning requires a shared update that makes balanced progress across multiple objectives. A key obstacle is that progress is difficult to compare across tasks with different loss scales, units, and local optimization dynamics. We study this problem from an optimality-gap perspective. For each task, we compare the improvement induced by the shared update with the best one-step improvement achievable by optimizing that task in isolation. This comparison yields a scale-invariant normalized improvement rate, which measures how well the shared update serves each task relative to its own local improvement bound. We formulate balanced multi-task optimization as a max-min problem over the normalized improvement rates, thereby prioritizing the worst-served task. A local quadratic approximation leads to the Karush-Kuhn-Tucker (KKT) optimality conditions, showing that the optimal shared update is determined by a small set of bottleneck tasks. Under an isotropic Hessian approximation, the update admits a closed-form expression that depends only on task gradients and a regularized Gram matrix, while the bottleneck set is identified by an active-set procedure. Experiments on diverse benchmarks show improved task balance and competitive performance.