Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Optimising Contextual Embeddings for Meaning Conflation Deficiency Resolution in Low-Resourced Languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
MosSunHla
Éditeur:
MDP
Hôte:
Meaning conflation deficiency (MCD) presents a continual obstacle in natural language processing (NLP), especially for low-resourced and morphologically complex languages, where polysemy and contextual ambiguity diminish model precision in word sense disambiguation (WSD) tasks. This paper examines the optimisation of contextual embedding models, namely XLNet, ELMo, BART, and their improved variations, to tackle MCD in linguistic settings. Utilising Sesotho sa Leboa as a case study, researchers devised an enhanced XLNet architecture with specific hyperparameter optimisation, dynamic padding, early termination, and class-balanced training. Comparative assessments reveal that the optimised XLNet attains an accuracy of 91% and exhibits balanced precision–recall metrics of 92% and 91%, respectively, surpassing both its baseline counterpart and competing models. Optimised ELMo attained the greatest overall metrics (accuracy: 92%, F1-score: 96%), whilst optimised BART demonstrated significant accuracy improvements (96%) despite a reduced recall. The results demonstrate that fine-tuning contextual embeddings using MCD-specific methodologies significantly improves semantic disambiguation for under-represented languages. This study offers a scalable and flexible optimisation approach suitable for additional low-resource language contexts.

Visit

doi.org

Languages

Sotho, NorthernSotho, Southern

Licenses

https://creativecommons.org/licenses/by/4.0/

Similaires

A Context-Aware Embedding Approach to Meaning Conflation Deficiency in Sesotho sa Leboa: Addressing Semantic AmbiguityMassive vs. Curated Embeddings for Low-Resourced Languages: the Case of Yorùbá and TwiMassive vs. Curated Word Embeddings for Low-Resourced Languages. The Case of Yorùbá and TwiContextual Text Embeddings for TwiSurface Realization Architecture for Low-resourced African LanguagesIsomorphic Cross-lingual Embeddings for Low-Resource Languages

A Context-Aware Embedding Approach to Meaning Conflation Deficiency in Sesotho sa Leboa: Addressing Semantic Ambiguity

A major problem in Natural Language Processing (NLP) is Meaning Conflation Deficiency (MCD), especia

Massive vs. Curated Embeddings for Low-Resourced Languages: the Case of Yorùbá and Twi

The success of several architectures to learn semantic representations from unannotated text and the availability of these kind of texts in online multilingual resources such as Wikipedia has facilitated the massive and automatic creation of resources for multiple

Massive vs. Curated Word Embeddings for Low-Resourced Languages. The Case of Yorùbá and Twi

The success of several architectures to learn semantic representations from unannotated text and the

Contextual Text Embeddings for Twi

Transformer-based language models have been changing the modern Natural Language Processing (NLP) la

Surface Realization Architecture for Low-resourced African Languages

There has been growing interest in building surface realization systems to support the automatic gen

Isomorphic Cross-lingual Embeddings for Low-Resource Languages

Cross-Lingual Word Embeddings (CLWEs) are a key component to transfer linguistic information learnt