Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

A Context-Aware Embedding Approach to Meaning Conflation Deficiency in Sesotho sa Leboa: Addressing Semantic Ambiguity

Domaine:

natural language processing

Type de record:

paper
Créateur:
MosSunHla
Éditeur:
Bil
Hôte:
A major problem in Natural Language Processing (NLP) is Meaning Conflation Deficiency (MCD), especially in low-resource, morphologically rich languages like Sesotho sa Leboa. In downstream tasks like Word Sense Disambiguation (WSD), traditional word embeddings frequently perform poorly because they are unable to distinguish between a word's numerous senses. To ascertain how well various context-aware and multi-prototype word embedding models—such as ELMo, GPT-2, BERT, Universal Sentence Encoder, and hybrid versions of Doc2Vec and SBERT—resolve MCD, this study examines and assesses them. Standard classification measures (precision, recall, F1-score, and accuracy) as well as clustering-based metrics and visualisation approaches were used to assess the models after they were trained and tested on a sense-annotated Sesotho sa Leboa corpus. According to the results, deep contextual models—in particular, ELMo and GPT-2—perform noticeably better in terms of accuracy and sense separation than static and unsupervised models. With well-separated confusion matrices, ELMo showed excellent interpretability and the highest F1-score (93%) of any model. According to the results, context-aware architecture provides reliable MCD solutions as well as a scalable framework for improving WSD in language applications with limited resources. For future studies on semantic disambiguation in under-represented languages, the work offers fresh standards and perspectives.

Visit

doi.org

Tasks

embeddings

Languages

Sotho, NorthernSotho, Southern

Licenses

https://creativecommons.org/licenses/by-nc/4.0

Similaires

A context-aware word embedding model for morphologically rich languages using Sesotho sa Leboa as a case studySesotho sa Leboa Spelling Checker 1.1Sesotho sa Leboa Genre Classification CorpusAfrican Wordnet: Sesotho sa Leboa 1.0Autshumato English-Sesotho sa Leboa Translation MemoryAutshumato English-Sesotho sa Leboa Parallel Corpora

A context-aware word embedding model for morphologically rich languages using Sesotho sa Leboa as a case study

Meaning conflation deficiency (MCD) is a major issue in natural language processing (NLP) to improve

Sesotho sa Leboa Spelling Checker 1.1

Spelling checkers and hyphenators for South African languages compatible with Microsoft® Office 2000

Sesotho sa Leboa Genre Classification Corpus

Contains training and testing data for Genre Classification for Sesotho sa Leboa.

African Wordnet: Sesotho sa Leboa 1.0

Developed using the expand model with Princeton WordNet 2.0 as basis. Each wordnet contains synsets

Autshumato English-Sesotho sa Leboa Translation Memory

Translation memory from English (EN-GB) to Sesotho sa Leboa, in the government domain for use in the

Autshumato English-Sesotho sa Leboa Parallel Corpora

Parallel corpora aligned on sentence level through a combination of automatic and manual alignment t