Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Enhancing Hausa Words Lemmatization Through Feature Engineering

Domaine:

natural language processing

Type de record:

datasetpaper
Créateur:
AdaRasMuh
Éditeur:
Fed
Hôte:
Hausa is spoken by over 50 million people across Africa, yet it remains critically under-resourced in natural language processing (NLP), particularly for lemmatization. The language’s rich morphology characterized by agglutination, internal vowel alternation, and extensive affixation poses significant challenges for existing rule-based and conventional machine learning approaches. This study addresses this gap by developing and evaluating supervised machine learning models for Hausa word lemmatization. We constructed a manually annotated dataset comprising 4,530 unique word-lemma pairs extracted from diverse media sources, achieving a high Inter-Annotator Agreement of 91.10%. Two baseline algorithms, Support Vector Machine (SVM) and Random Forest, were trained and optimized using GridSearchCV on an 80/20 train-test split. The study introduces an enhanced feature engineering framework that integrates phonological attributes, morphological markers, syllable counts, gemination flags, and extended character n-grams (1-5) alongside traditional surface-level features. Experimental results demonstrate that the Random Forest classifier consistently outperforms SVM. When paired with the enhanced feature set, Random Forest achieved the highest performance metrics, recording an accuracy of 64.02% and a weighted F1-score of 0.6134. Feature importance analysis further confirms that linguistically informed attributes significantly improve model generalization and prediction accuracy. These findings underscore the critical role of domain specific feature design in overcoming data scarcity and linguistic complexity. The curated dataset and optimized modeling framework provide a foundational resource to advance downstream Hausa NLP applications, including information retrieval, machine translation, and computational linguistics.

Visit

doi.org

Languages

Hausa

Licenses

https://creativecommons.org/licenses/by/4.0