Logo Lanfrica

Curation of a polysemous word dataset for word sense disambiguation in Hausa language

Domaine:

natural language processing

Type de record:

dataset
Créateur:
HalI. P.O
Éditeur:
Jou
Hôte:
The challenge of Word Sense Disambiguation (WSD) is fundamental to Natural Language Processing (NLP), particularly in low-resource languages where lexical ambiguity hinders effective language understanding. Hausa, a major Chadic language spoken by over 60 million people, lacks structured lexical resources for disambiguating polysemous words. This paper presents the development and curation of a high-quality Hausa Polysemous Word Sense Disambiguation dataset consisting of 2,021 manually selected and annotated lemmas. Each lemma is disambiguated into its distinct senses, accompanied by contextual Hausa example sentences, English glosses, and translations. The dataset is designed to support the training and evaluation of supervised and semi-supervised WSD models for Hausa and serves as a foundational resource for semantic NLP tasks in low-resource settings. The annotation schema, curation methodology, and linguistic validation process are described in detail. This work fills a critical gap in Hausa NLP and provides a reproducible framework for constructing sense-annotated corpora in other under-resourced languages.