Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

YoNER: A New Yorùbá Multi-domain Named Entity Recognition Dataset

Domaine:

natural language processing

Type de record:

paperdatasetmodel
Créateur:
FalAlaAkiOgu
Hôte:avatar
Named Entity Recognition (NER) is a foundational NLP task, yet research in Yorùbá has been constrained by limited and domain-specific resources. Existing resources, such as MasakhaNER (a manually annotated news-domain corpus) and WikiAnn (automatically created from Wikipedia), are valuable but restricted in domain coverage. To address this gap, we present YoNER, a new multidomain Yorùbá NER dataset that extends entity coverage beyond news and Wikipedia. The dataset comprises about 5,000 sentences and 100,000 tokens collected from five domains including Bible, Blogs, Movies, Radio broadcast and Wikipedia, and annotated with three entity types: Person (PER), Organization (ORG) and Location (LOC), following CoNLL-style guidelines. Annotation was conducted manually by three native Yorùbá speakers, with an inter-annotator agreement of over 0.70, ensuring high quality and consistency. We benchmark several transformer encoder models using cross-domain experiments with MasakhaNER 2.0, and we also assess the effect of few-shot in-domain data using YoNER and cross-lingual setups with English datasets. Our results show that African-centric models outperform general multilingual models for Yorùbá, but cross-domain performance drops substantially, particularly for blogs and movie domains. Furthermore, we observed that closely related formal domains, such as news and Wikipedia, transfer more effectively. In addition, we introduce a new Yorùbá-specific language model (OyoBERT) that outperforms multilingual models in in-domain evaluation. We publicly release the YoNER dataset and pretrained OyoBERT models to support future research on Yorùbá natural language processing. LREC 2026

Visit

arxiv.org

Tasks

named entity recognitioninformation extraction

Languages

Yoruba

Tags

Computation and Language

Similaires

Robust Yorùbá Named Entity Recognition through Simple Mixed TrainingDzNER: A large Algerian Named Entity Recognition datasetGovernment Domain Named Entity Recognition for South African LanguagesAn Open-Source Dataset and A Multi-Task Model for Malay Named Entity RecognitionNERAMazigh: A Named Entity Recognition Dataset for the Amazigh LanguageAMHARIC NAMED ENTITY RECOGNITION IN THE AGRICULTURE DOMAIN USING A MACHINE LEARNING APPROACH

Robust Yorùbá Named Entity Recognition through Simple Mixed Training

Yorùbá named entity recognition (NER) is sensitive to missing tone marks and in-line English, both c

DzNER: A large Algerian Named Entity Recognition dataset

Government Domain Named Entity Recognition for South African Languages

This paper describes the named entity language resources developed as part of a development project for the South African languages. The development efforts focused on creating protocols and annotated data sets with at least 15,000 annotated named entity tokens for

An Open-Source Dataset and A Multi-Task Model for Malay Named Entity Recognition

Named entity recognition (NER) is a fundamental task of natural language processing (NLP). However,

NERAMazigh: A Named Entity Recognition Dataset for the Amazigh Language

NERAMazigh is a manually annotated Named Entity Recognition (NER) dataset for the Amazigh language,

AMHARIC NAMED ENTITY RECOGNITION IN THE AGRICULTURE DOMAIN USING A MACHINE LEARNING APPROACH