YoNER: A New Multi-Domain Named Entity Recognition Dataset and Monolingual Pretrained Language Model for Yorùbá
Abstract
Named entity recognition (NER) underpins information extraction, but while high-resource languages have large, multi-domain corpora, low-resource languages like Yorùbá (spoken by 50+ million people) have NER resources confined to news text, limiting model performance on other genres. This work introduces YoNER, the first multi-domain Yorùbá NER dataset, to study: (1) how news-trained NER models transfer across domains (cross-domain transfer); (2) whether small in-domain data improves transfer; and (3) whether a monolingual Yorùbá pretrained language model (PLM) outperforms multilingual PLMs for multi-domain transfer. YoNER spans five domains: Bible, Blogs, Movies, Radio, and Wikipedia, totalling 5,148 sentences (100,795 tokens), annotated by three native speakers under CoNLL guidelines. MasakhaNER 2.0 (Adelani et al., 2022) was included for the news domain. We also pretrained OyoBERT, a Yorùbá pretrained language model using a token-dropping objective with some Yorùbá subset datasets and machine-translated English data, which produced base and large variants configured analogously to BERT-base and BERT-large, with an ablation variant (YoBERT) excluding the translated data. OyoBERT was compared with multilingual PLM, AfroXLMR. Generally, machine-translated pretraining data improved OyoBERT-base over YoBERT-base by at least three F1 points, and scaling to OyoBERT-large yielded further gains, achieving the highest average score among monolingual models evaluated. For cross-domain transfer, AfroXLMR-large-76L led overall, though Blogs and Movies remained hardest. Small amounts of in-domain data substantially improved transfer: Bible achieved the highest in-domain performance (84.9%), and, alongside News, Wikipedia and Radio, proved to be the most effective source and target domains overall. Monolingual OyoBERT-large excelled in in-domain transfer, while AfroXLMR-large led under multi-domain transfer. YoNER and OyoBERT together offer reusable resources advancing domain-diverse, low-resource African NLP. YoNER contributes the first multi-domain, native-speaker-validated Yorùbá NER dataset, alongside OyoBERT, a new monolingual Yorùbá PLM trained with synthetic data augmentation.