Delve into the Applicability of Advanced Optimizers for Multi-Task Learning
Abstract
Multi-Task Learning (MTL) is a fundamental problem in machine learning that has been extensively studied over the past decade. Recently, a variety of optimization-based MTL approaches have been proposed to jointly learn multiple tasks by modifying the optimization trajectory. In this paper, we argue that the design of mainstream optimization-based MTL methods implicitly assumes compatibility with momentum-free optimizers, which may limit the understanding of their effectiveness in modern training regimes. In practice, the instantaneously derived gradients from MTL operations contribute only marginally to the final parameter updates when momentum is present, resulting in an amortized rather than instant de-conflicting effect. Moreover, this amortized behavior can be further degraded under high-curvature optimization dynamics. To address this issue and enable the effective integration of mainstream MTL methods with advanced optimizers (e.g., Adam and Muon), we propose \texttt{APT} (Applicability of advanced oPTimizers), a lightweight framework featuring an adaptive momentum mechanism that balances the trade-off between instant and amortized de-conflicting. Furthermore, we show that the Muon optimizer can be interpreted as an implicit MTL learner, and we introduce a lightweight direction preservation strategy to better align with its orthogonalization process. Extensive experiments across four standard MTL benchmarks demonstrate that \texttt{APT} consistently enhances existing MTL approaches, yielding substantial performance gains.