Moving Past Random Splits: Audited Evaluation and Parameter-Efficient Methods for Low-Resource South African Sign Language Recognition
Abstract
Machine learning for sign language recognition has progressed quickly, but advances for under-resourced sign languages are still limited by the scarcity of annotated data and by evaluation protocols where duplicate samples or overlap in signers/sessions can inflate perceived generalisation performance. We address this problem in the context of 27-class static South African Sign Language (SASL) alphabet recognition, placing emphasis on both rigorous evaluation and parameter-efficient representation learning. The study relies on three pre-existing South African static-image datasets instead of collecting new participant data. In total, 49,586 repository images were screened for exact and near duplicates, as well as for blur and file corruption. Since the original metadata does not include confirmed signer identities, we do not present our evaluation as fully signer-independent. Rather, we construct the benchmark from the subset with the most reliable metadata for deriving proxy signer/session groupings, and we partition at the level of these groups instead of individual images. The final split comprises 10,872 training, 1,919 validation, and 4,323 held-out test images over 27 classes (A–Z plus SPACE), with overlapping proxy groups segregated and no exact-duplicate leakage across splits. Under identical preprocessing, optimization settings, random seeds, loss weighting, and batch-ordering, we evaluate six learning approaches: a task-specific CNN, CNN-BiLSTM, CNN-LSTM, a CNN-GAP ablation, a frozen MobileNetV2 transfer baseline, and a linear SVM operating on CNN-extracted features. The CNN-BiLSTM attains 99.93% test accuracy and 99.93% macro F1 with 363,675 parameters, coming close to the CNN’s perfect 100.00% accuracy while using roughly 71.6% fewer parameters than the 1.28M-parameter CNN. By comparison, the frozen MobileNetV2 and linear SVM baselines yield 72.26% and 71.06% accuracy, respectively. Further quantitative error analysis indicates that models with weaker representations tend to concentrate their errors on visually similar hand configurations. These results show that, on this carefully controlled low-resource SASL benchmark, learning task-specific representations is markedly more effective than relying on fixed, generic visual features, and that spatial-sequential aggregation offers a parameter-efficient substitute for a deeper CNN. More generally, the findings underscore the need for leakage-aware, signer-sensitive evaluation when interpreting the performance of advanced sign-language recognition systems. Future work will expand this benchmark to support rigorously verified signer-independent evaluation and continuous SASL recognition directly from video.