Concept-Aware Wasserstein Routing with Vision-Language Guidance for Few-Shot WSI Classification
Abstract
Few-shot whole-slide image (WSI) classification remains challenging due to gigapixel-scale images, scarce slide-level labels, and strong morphological heterogeneity. Recent vision-language MIL methods improve data efficiency by transferring knowledge from pathology foundation models, but they often treat patches as weakly related instances and represent each class with a single semantic or visual anchor. We propose a multi-scale vision-language framework that combines prompted pathology encoding, pathology concept learning, Semantic Wasserstein Routing (SWR), and barycentric prototype memory. Using frozen pathology vision-language backbones with lightweight prompt adaptation, our method aligns image patches with concept embeddings through unbalanced optimal transport and builds a semantic patch graph for concept-guided aggregation. In parallel, the prototype memory maintains multiple class-specific Wasserstein barycenters to capture diverse morphological modes. A distance-based fusion strategy integrates slide-level text alignment, concept-guided patch evidence, and prototype-memory matching for robust prediction. Experiments on TCGA WSI classification benchmarks under few-shot settings show consistent improvements over strong MIL and vision-language MIL baselines, demonstrating the benefit of semantic patch routing and multi-prototype reasoning for weakly supervised pathology learning.