WaveMamba: Wave-Inspired Cross-Modal Fusion for Robust event-image Semantic Segmentation
Abstract
Event-image semantic segmentation is promising for autonomous driving and robotics because event cameras complement frame images under rapid motion, low illumination, and high dynamic range. However, effective multimodal segmentation remains difficult due to asynchronous sensing, sparse-dense modality mismatch, and the cost of heavy cross-modal interaction modules. We propose WaveMamba, a dual-branch event-image semantic segmentation framework that combines hierarchical GroupMamba encoders with a Mamba-Wave Cross-Modal Fusion (MWCMF) module. Instead of relying on explicit registration or attention-heavy fusion, MWCMF performs implicit cross-modal propagation through channel-wise and spatial-wise wave-guided gating in the spectral domain, followed by Mamba-based refinement. Experiments on DDD17 and DSEC show that WaveMamba achieves state-of-the-art 78.75 and 76.00 mIoU, respectively. Under rain- and fog-corrupted evaluation on DSEC, WaveMamba further attains 74.0 mIoU, indicating improved robustness under degraded visual conditions.