From Clips to Streams: A Unified Framework for Streaming Sign Language Translation
Abstract
Sign Language Translation (SLT) research has typically focused on pre-segmented clips, a constraint that disconnects models from the continuous reality of real-world communication. To bridge the gap from clips to streams, we formally define \textbf{Streaming SLT}, a realistic task demanding the simultaneous discovery and translation of linguistic events from continuous, untrimmed video inputs. To address this challenge, we propose \textbf{StreamSLST}, the first end-to-end framework for streaming \textbf{S}ign \textbf{L}anguage \textbf{S}egmentation and \textbf{T}ranslation via Deformable Transformer architecture. Unlike cascaded approaches that suffer from error propagation, StreamSLST employs a parallel decoding strategy to jointly optimize \textit{temporal segmentation} and \textit{translation} in a single pass. To enable scalable training on long-form videos, we incorporate a \textit{Sentence-aware Sliding Window Sampler} that crops continuous streams into trainable units. We further design a \textit{Tri-modal Visual-Language Pretraining} stage that aligns pose dynamics with textual semantics before joint optimization. To support comprehensive evaluation, we conduct experiments on two real-world streaming datasets, BOBSL and How2Sign, together with two newly constructed synthetic streaming benchmarks, \textit{Streaming-CSL-Daily} and \textit{Streaming-Phoenix-2014T}. The results demonstrate that StreamSLST consistently outperforms cascaded segmentation-then-translation pipelines. Interestingly, our analysis reveals that learned sentence boundaries can outperform ground-truth boundaries for translation, suggesting that joint optimization drives localization toward semantically informative regions rather than exact temporal endpoints. Our work establishes the first robust benchmark for continuous sign language understanding and paves the way for accessible real-world communication interfaces.