Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Data for the Journal Paper Machine Learning-Based Context-Aware Lemmatization for Low-Resource Languages: A Case Study of Setswana

Domaine:

natural language processing

Type de record:

datasetpapersoftware
Créateur:
Iri
Hôte:avatar

Efficient natural language processing (NLP) tools for Setswana are essential for improving human-machine interaction, yet the language remains underrepresented in computational linguistics due to its complex morphology and limited linguistic resources. This study introduces a context-aware machine-learning-based lemmatization model for Setswana, addressing challenges in word sense disambiguation and morphological analysis. Unlike previous rule-based lemmatizers, which process words in isolation, this model incorporates contextual information using Naïve Bayes (NB) and N-gram embeddings to improve lemma prediction accuracy. The proposed model was trained and evaluated using a manually annotated Setswana corpus, integrating part-of-speech (POS) tagging and named entity recognition (NER) as key linguistic features. Performance evaluation, based on accuracy (70.32%), precision (70%), recall (65%), and F1-score (66%), demonstrates the model’s effectiveness in resolving polysemous words, a challenge not addressed by existing Setswana lemmatization approaches. Comparative analysis with prior studies highlights that machine-learning models outperform rule-based approaches in capturing contextual dependencies, although dataset size and feature selection remain critical to performance improvement. This research marks a significant advancement in Setswana NLP, establishing a foundation for future hybrid models that integrate deep learning and rule-based techniques for enhanced accuracy. The study contributes to the development of computational tools for low-resource languages, paving the way for their inclusion in modern information retrieval, machine translation, and conversational AI systems.

Visit

figshare.com

Languages

Setswana

Tags

Machine learning not elsewhere classifiedAfrican languagesComputational linguisticsLemmatizationSetswanamachine learningnaive Bayesian classification (NBC)context-aware processingNatural Language Processing Model

Licenses

CC BY 4.0