Auditing the Evaluation of Sign Language Production with Content-Free Baselines
Abstract
Sign language is a dense temporal signal which carries meaning through the coordinated movement of hands, face and body. This makes the performance of sign language production (SLP) models difficult to evaluate. In the literature, it is typically measured in two ways: a distance between the generated motion and a reference sequence, and the ability of a sign language translation model to recover the intended sentence from the generated motion, known as back-translation. Both numbers are reported and compared as absolute scores, yet neither is anchored: nothing in the protocol tells the reader which values indicate learnt signing and which indicate none. We audit what these numbers can and cannot tell us about a model's ability to sign. Using the selected content-free baseline signals, we measure what the current protocol awards to outputs that carry no linguistic content, and find that the gains reported by recent approaches are difficult to interpret. We investigate two state-of-the-art SLP models and illustrate how a model can score well on both metrics without having learnt to sign. We propose a protocol that grounds them against measured floors and ceilings and a shared translation model, turning quantitative evaluation into a more concrete evidence of what an SLP model has actually learnt.