Don't Pay Attention, PLANT It: Pretraining Attention via Learning-to-Rank
Abstract
State-of-the-art Extreme Multi-Label Text Classification (XMC) models rely on multi-label attention to focus on key tokens in input text, but learning high-quality attention weights is challenging. We introduce PLANT (Pretrained and Leveraged Attention), a plug-and-play strategy for initializing attention. PLANT works by \emph{planting} label-specific attention using a pretrained Learning-to-Rank model guided by mutual information gain. This architecture-agnostic approach integrates seamlessly with large language model backbones such as Mistral, LLaMA, DeepSeek, and Phi-3. PLANT outperforms state-of-the-art methods across tasks such as ICD coding, legal topic classification, and content recommendation. Gains are especially pronounced in few-shot settings, with substantial improvements on rare labels. Ablation studies confirm that attention initialization is a key driver of these gains. Code and trained models are available at \url{https://github.com/research-anon-487/xcube/tree/plant}.