Enhancing Clinical Interpretability in Whole Slide Imaging via Pathology Foundation Models and Attribute-based MIL
Fatemeh Hadizadeh ⋅ Mahsa Jalali
Abstract
Digital Pathology and Whole Slide Imaging (WSI) present significant computational challenges due to their gigapixel resolution. While Multiple Instance Learning (MIL) is the standard paradigm, current frameworks often rely on generic image encoders (e.g., ResNet-50), which act as black boxes and fail to capture complex pathological morphologies. In this work, we integrated state-of-the-art Pathology Foundation Models (specifically UNI2-h) with the AttriMIL framework to classify non-small cell lung cancer (TCGA-NSCLC). We successfully engineered a data harmonization pipeline to project high-dimensional semantic embeddings ($D=1536$) into the framework's native latent space, while independently constructing a structured $k$-Nearest Neighbor ($k$-NN) spatial graph capturing physical tissue topology. Operating under severe hardware and infrastructure limitations, our optimized pipeline demonstrated unprecedented data efficiency and accelerated training dynamics, converging to an optimal state in just 18 epochs using only half the standard training data. Evaluated on a strictly isolated patient-level hold-out set of unseen WSIs via Monte Carlo bootstrapping, our approach achieved a state-of-the-art AUC-ROC of $0.9742 \pm 0.0063$ and an Accuracy of $0.9220 \pm 0.0116$. Furthermore, we developed a comprehensive Vision-Language Clinical Alignment strategy utilizing NLP and Point-Biserial correlation, complemented by a qualitative morphological validation pipeline for clinical hallmarks like keratinization and necrosis, proving the clinical interpretability of our model.
Chat is not available.
Successful Page Load