Contrastive Retrieval Heads for Improved Attention-Based Reranking
Abstract
The strong zero-shot and long-context capabilities of recent Large Language Models (LLMs) have enabled highly effective list-wise reranking systems. Attention-based rerankers leverage transformer attention patterns as retrieval signals, but not all attention heads contribute equally: many introduce noise and redundancy, limiting both reranking quality and efficiency. In this work, we introduce CoRe heads, a small subset of retrieval-specialized heads identified through a contrastive scoring objective that rewards attention to relevant documents while penalizing attention to distractors. This contrastive criterion isolates highly discriminative retrieval heads and yields a strong training-free list-wise reranker. Our analysis reveals that retrieval-relevant computation in transformer rerankers is highly sparse and structurally localized: fewer than 1% of attention heads account for most reranking performance across models and datasets, and these heads consistently concentrate in middle transformer layers. This localization enables aggressive layer pruning for efficient reranking: removing the final 50% of transformer layers preserves nearly identical reranking accuracy while substantially reducing inference latency and GPU memory usage.