Lost in Translation, Found in Embeddings: Sign Language Translation and Alignment
Abstract
Our aim is to develop a unified model for sign language understanding that performs both sign language translation (SLT) and sign–subtitle alignment (SSA). Together, these two tasks enable the conversion of continuous signing videos into spoken language text and also the temporal alignment of signing with subtitles -- both beneficial for practical communication, large-scale corpus construction, and educational applications. To achieve this, our approach is built upon three design choices: (i) a lightweight visual backbone that captures manual and non-manual cues from human keypoints and lip-region images while preserving signer privacy and enabling end-to-end training; (ii) a Sliding Perceiver mapping network that aggregates consecutive visual features into word-level embeddings to bridge the vision–text gap; and (iii) scaling training to large datasets -- BOBSL covering British Sign Language (BSL) and YouTube-SL-25 covering American Sign Language (ASL) -- to promote generalisation across languages and signers. With this multi-task, multilingual pretraining and strong model design, we achieve state-of-the-art results on BOBSL, How2Sign, Phoenix14T, and FLEURS-ASL SLT benchmarks, and also for SSA on BOBSL and WMT-SLT-SRF. Beyond these standard, signer-interpreted datasets, we show successful SLT on the natural signing data from the BSL-Corpus.