Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Data for the Journal Paper Machine Learning-Based Context-Aware Lemmatization for Low-Resource Languages: A Case Study of Setswana

Domain:

natural language processing

Record type:

datasetpapersoftware
Creator:
Iri
Host:avatar

Efficient natural language processing (NLP) tools for Setswana are essential for improving human-machine interaction, yet the language remains underrepresented in computational linguistics due to its complex morphology and limited linguistic resources. This study introduces a context-aware machine-learning-based lemmatization model for Setswana, addressing challenges in word sense disambiguation and morphological analysis. Unlike previous rule-based lemmatizers, which process words in isolation, this model incorporates contextual information using Naïve Bayes (NB) and N-gram embeddings to improve lemma prediction accuracy. The proposed model was trained and evaluated using a manually annotated Setswana corpus, integrating part-of-speech (POS) tagging and named entity recognition (NER) as key linguistic features. Performance evaluation, based on accuracy (70.32%), precision (70%), recall (65%), and F1-score (66%), demonstrates the model’s effectiveness in resolving polysemous words, a challenge not addressed by existing Setswana lemmatization approaches. Comparative analysis with prior studies highlights that machine-learning models outperform rule-based approaches in capturing contextual dependencies, although dataset size and feature selection remain critical to performance improvement. This research marks a significant advancement in Setswana NLP, establishing a foundation for future hybrid models that integrate deep learning and rule-based techniques for enhanced accuracy. The study contributes to the development of computational tools for low-resource languages, paving the way for their inclusion in modern information retrieval, machine translation, and conversational AI systems.

Visit

figshare.com

Languages

Setswana

Tags

Machine learning not elsewhere classifiedAfrican languagesComputational linguisticsLemmatizationSetswanamachine learningnaive Bayesian classification (NBC)context-aware processingNatural Language Processing Model

Licenses

CC BY 4.0