Attention Is Not the Only Readout: A Free Localization Signal in Trained MIL Models, and the Regime Where It Fails
Abstract
A weakly supervised multiple-instance learning (MIL) model is trained only on slide-level labels, yet is routinely asked which regions of the slide carry the evidence. The field's standard answer is the attention map. We show that it is only one of several per-patch maps already contained in a trained model, that it is usually the weakest of them, and that, across most of our experimental grid, the choice of readout changes measured localization more than the choice of architecture. Three maps can be extracted from identical trained weights with no retraining, no added parameters, and no annotations. The first is the attention weight assigned to each patch. The second, proposed in prior work, is each patch's attention-weighted contribution to the slide prediction. The third, which we advocate, applies the model's own slide classifier independently to each patch embedding and requires only a single matrix multiplication. We compare all three under a unified protocol across two pathology foundation-model encoders (UNI and Virchow2, extracted on identical patch grids), two MIL architectures (ABMIL and CLAM-SB), and three cohorts spanning two organs, scoring every map against public pixel-level annotations. Each configuration is trained with five random seeds (three on PANDA), and readouts are compared using paired bootstrap and Wilcoxon tests over slides. On Camelyon16, reading ABMIL through the classifier rather than attention raises the per-slide patch AUC from 0.847 to 0.978 (+0.131, 95% CI [+0.099, +0.168]), outperforming attention on 96% of slides and reducing the proportion of slides localized below 0.9 AUC from 58% to 6%. The same models applied to Camelyon17, five centres removed from training, show the largest gains on the smallest and most difficult lesions: isolated tumour cells improve from 0.881 to 0.986 (across the 11 of 16 scorable ITC slides), micrometastases from 0.841 to 0.961, compared with only +0.047 on macrometastases. Replacing ABMIL with CLAM-SB while keeping the readout fixed changes the same metric 1.7 to 4.0 times less. The gap is not an artefact of checkpoint selection. Evaluated across eleven training checkpoints, attention localization peaks after a single epoch (0.946) and then declines monotonically to 0.850, whereas the classifier readout remains stable throughout training (0.954–0.979). Consequently, the performance gap grows from +0.011 to +0.127 as the slide-level objective converges, and at no checkpoint does attention become the better readout. The advantage is not universal. Across 1,187 annotated prostate biopsies, the ordering reverses for ABMIL, with attention outperforming the classifier (0.832 vs. 0.756), while the classifier remains superior for CLAM-SB (0.822 vs. 0.791). To compare cohorts of differing difficulty on a common scale, we train a supervised patch classifier using training-slide annotations and express every readout as its chance-corrected recovery of that reference. Attention remains within a narrow 14-point band (68–81%) across every cohort, encoder, and architecture, whereas the classifier readout spans 40 points, ranging from 95–99% on lymph node datasets to 60–75% on prostate. This difference characterizes how the regimes differ rather than explaining why. Finally, we show that two cross-slide normalization conventions, both defensible and neither typically reported, reverse the lesion-level FROC ranking in three of four experimental settings despite using identical model checkpoints. We recommend that localization studies explicitly report the readout used, include at least one alternative readout, and benchmark both against a supervised reference.