SICAF: Time-Varying Focus Bottleneck for Self-Supervised Event-Based Optical Flow with Spiking Neural Network
Shuangming Yang ⋅ Qing He ⋅ Shangqi Guo ⋅ Badong Chen
Abstract
Event-based optical flow estimation achieves high temporal resolution and dynamic range by exploiting asynchronous events, while self-supervised learning provides an effective solution to the absence of ground-truth labels. Spiking neural networks (SNNs), as brain-inspired models, process information through asynchronous spikes, naturally aligning with optical flow events. However, the temporal irregularity across event streams hinders the compact extraction of spatiotemporal features, while the absence of explicit labels leads to spurious correlations, thereby increasing redundancy within latent representations. In this paper, we propose a novel “Compactly Compress, Time-varying Focus” training strategy. We present SICAF ($\text{\textbf{S}pikes}$ $\text{\textbf{I}nformation}$ $\text{\textbf{C}ompression}$ $\text{\textbf{A}nd}$ $\text{\textbf{F}ocus}$) framework that formulates the learning process as a constrained optimization with temporal modulation. It encourages self-supervised SNNs to extract compact latent representations while preventing over-compression that would discard temporal saliency. To better enable feature selectivity across time steps, we design an estimator with learnable temporal attention, achieving context awareness with cross-time dependencies. Experimental results on several datasets demonstrate that SICAF achieves state-of-the-art performance in self-supervised SNNs, with a 8.14\% improvement in AEE on the MVSEC dataset and a 13.49\% increase in robustness under white-box adversarial attacks.
Chat is not available.
Successful Page Load