LSC-Parlament: An Automatically Aligned Catalan Sign Language Dataset from Parliament Videos.
Abstract
Progress in Sign Language Translation (SLT) is frequently hampered by the "data bottleneck," a challenge particularly acute for regional languages such as Catalan Sign Language (LSC). In this paper, we present LSC-Parlament, a large-scale, multimodal dataset for LSC derived from Catalan Parliament sessions spanning 2021 to 2024. The dataset comprises over 87 hours of LSC video, synchronized with speech and text translations in either Catalan or Spanish. To curate this resource, we developed a fully automated pipeline that enables alignment between sign language, speech, and text without requiring prior human-annotated data. We provide a systematic evaluation of the dataset across multiple dimensions, including video tracking quality, audio transcription accuracy, and translation benchmarks. Our experiments compare transfer learning, zero-shot, and supervised translation settings, demonstrating significant knowledge transfer between spoken languages in the LSC-Spanish modality. By releasing LSC-Parlament, we provide a challenging open-domain benchmark to foster research in inclusive and scalable translation technologies.