Topology-Reinforced Swin Transformer for Medical Image Analysis
Pengfei Gu ⋅ Huimin Li ⋅ Guangyu Meng ⋅ Hao Zheng ⋅ Haoteng Tang ⋅ Bin Fu ⋅ Danny Z Chen
Abstract
Topological patterns in medical images, such as connected components and loops, across multiple spatial scales carry critical structural information, from microaneurysm-scale lesions to organ-level boundaries. Despite the success of hierarchical Vision Transformers (e.g., Swin Transformer) in medical image analysis, existing methods lack an explicit architectural scheme to represent and incorporate multi-scale topological structures through the Transformer hierarchy. In this paper, we propose \emph{Topology-Reinforced Swin Transformer (TRiST)}, a new Swin Transformer model family for medical image analysis, which reinforces multi-scale hierarchical topology into the hierarchical Transformer architecture. Instead of treating topology as an auxiliary descriptor or post hoc fusion cues, TRiST takes topology as a native part of representation construction, refinement, and token interaction. First, because topology does not compose hierarchically through pooling or interpolation, we develop a Hierarchical Multi-Scale Topology Representation algorithm that computes 2D persistent homology ($H_0$ and $H_1$ on image patches) at each native scale of Swin Transformer, constructing four topological representations aligned 1-to-1 with the Swin token grid. Second, since the persistence axis has an intrinsic semantic structure (i.e., short-lifetime features mainly capture noise and fine texture, whereas long-lifetime features capture stable anatomical structures), we design a Band-Aware $H_0/H_1$ Residual Refinement module that independently adapts six persistence sub-bands using dedicated gated residual MLPs. Third, to make topology take part in token interaction inside the Transformer, we introduce a Persistence-Band Progressive Attention Bias that injects stage-corresponding topology as an additive key-highlighting prior into shifted-window self-attention. Based on these strategies, we instantiate two task-specific models: Topo-SwinV2-B for classification and Topo-Swin-UNet for segmentation. Both models retain the original Swin-based backbone structure while equipping it with native multi-scale topological features. Experiments on five medical image datasets show consistent improvements over strong Swin-based baselines and competitive performance as existing topology-augmented methods.
Chat is not available.
Successful Page Load