Quantifying MIL Prediction Robustness and Instability in Histology WSI: Towards Responsible AI
Abstract
Multiple-instances learning (MIL) has become a widely used approach to whole-slide image (WSI) classification because it learns from slide level labels only. However, responsible deployment requires more than discriminative accuracy: a model should give consistent predictions for parallel sections of the same block, and standard measures such as the area under the ROC curve (AUC) do not capture this. We propose a quantitative measure of the robustness of MIL prediction that takes advantage of common practice in any pathology laboratory - multiple parallel cuts of the same paraffin block are on one or more slides. The corresponding cuts were identified with a YOLO-based detector, each cut was classified independently, and the instability was quantified as the range of predicted tumor probabilities within a parallel-cut group. We applied this measure to Clustering-constrained Attention MIL (CLAM) and Attention-based Deep MIL (ABMIL) for tumour versus non-tumour classification of head-and-neck biopsy WSIs across training sizes from 150 to 900 slides. Both models achieved test AUC above 0.92, and at 900 slides CLAM reached a marginally higher AUC than ABMIL (0.9611 versus 0.9539). Under the robustness measure the picture reversed: CLAM produced more frequent and larger deviations across parallel cuts, while ABMIL gave consistent probabilities. AUC alone would therefore favour the less stable model. Parallel-cut analysis quantifies a dimension of model behaviour that accuracy metrics miss, requires no additional annotation, and offers a practical route to identifying predictions that warrant pathologist review — a necessary step toward responsible AI in histopathology.