Event-SAMA: Language-Grounded Video Object Segmentation on Neuromorphic Event Streams
Sergei Fedorchenko ⋅ Max Zehnder ⋅ Nikola Zubić ⋅ Rong Zou ⋅ Davide Scaramuzza
Abstract
Event cameras provide microsecond temporal resolution, high dynamic range, and no motion blur, yet language-grounded video segmentation has been developed exclusively on dense RGB frames: no dataset and no model exist for referring video object segmentation (RVOS) on event streams. We address this gap with a dataset and a baseline model. First, we present **Event-SAMA-239K**, a synthetic-event conversion of the SAMA-239K grounded video-chat corpus and of all additional SAMA training data: 30,650 video clips and 392,162 images, converted with optical-flow, FILM, or saccade upsampling followed by a linear event generator, with all mask, referring-expression, and dialogue annotations preserved. Second, we present **Event-SAMA**, an adaptation of the SAMA-4B architecture that consumes events natively through a two-channel polarity patch embedding initialized from the pretrained RGB weights and a zero-initialized event adapter. Trained end-to-end on the converted corpus, Event-SAMA reaches $24.37$ J&F on MeViS `valid_u` and $22.33$ J&F on the converted subset of ReVOS `valid`, establishing the first event-native RVOS baseline (RGB SAMA-4B: $55.4$ on MeViS). A seed-controlled ablation with a measured noise floor of $0.47$ J&F identifies three design choices that matter: training on a diverse source mix instead of MeViS alone ($+4.99$), RGB-mean initialization of the event stem ($+4.33$), and a trainable SAM2 mask decoder ($+2.76$); the event adapter, partial ViT unfreezing, and event polarity are within noise at this budget. A comparison run with the identical recipe on an earlier, partially converted snapshot of the corpus scores $1.38$ J&F higher on MeViS and $0.73$ higher on ReVOS. Both are single runs that also differ in schedule length, so we draw no conclusion about corpus scale from this pair.
Chat is not available.
Successful Page Load