Foveal-Mamba: Inside-Out Ring Scanning with Recurrent Offset Prediction for Visual State Space Models
Abstract
Visual state space models (SSMs) have emerged as efficient alternatives to attention-based backbones, yet their performance remains sensitive to how visual evidence is sampled and propagated. Deformable visual Mamba methods address fixed scanning by predicting adaptive sampling locations, but offset prediction is still driven primarily by local responses, producing locally plausible samples that are weakly coupled to the SSM's recurrent state evolution, particularly in off-center or multi-object scenes. We present Foveal-Mamba, a scan-path-aware deformable visual Mamba framework that reformulates offset prediction as recurrent scan-path prediction. Two components drive this design. First, an Inside-Out Ring Scan organizes visual tokens along a center-to-periphery path, providing a stable propagation prior that does not assume object centering. Second, a lightweight Mamba module embedded within the offset prediction network forms the Recurrent Scan-Path Offset Prediction Network (RS-OPN), which predicts offsets from hidden states accumulated along the ring path. RS-OPN tightly couples deformable sampling with recurrent state evolution and produces coherent, object-aware sampling trajectories. Foveal-Mamba consistently outperforms strong baselines across multiple benchmarks. On ImageNet-1K, Foveal-Mamba-Tiny achieves 84.3\% Top-1 accuracy, surpassing DAMamba-Tiny by 0.5\% with fewer parameters and FLOPs. On ADE20K semantic segmentation, Tiny/Small variants reach 50.8/51.8 mIoU. On COCO, it achieves 48.9/50.2 box AP and 43.7/44.9 mask AP for object detection and instance segmentation, respectively. Qualitative results further demonstrate earlier object focus and more stable activations in multi-object scenes.