Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A Multilingual Sentiment Lexicon for Low-Resource Language Translation using Large Languages Models and Explainable AI

Domain:

natural language processing

Record type:

paperdatasetmodel
Creator:
MalLupNkovan
Host:avatar
South Africa and the Democratic Republic of Congo (DRC) present a complex linguistic landscape with languages such as Zulu, Sepedi, Afrikaans, French, English, and Tshiluba (Ciluba), which creates unique challenges for AI-driven translation and sentiment analysis systems due to a lack of accurately labeled data. This study seeks to address these challenges by developing a multilingual lexicon designed for French and Tshiluba, now expanded to include translations in English, Afrikaans, Sepedi, and Zulu. The lexicon enhances cultural relevance in sentiment classification by integrating language-specific sentiment scores. A comprehensive testing corpus is created to support translation and sentiment analysis tasks, with machine learning models such as Random Forest, Support Vector Machine (SVM), Decision Trees, and Gaussian Naive Bayes (GNB) trained to predict sentiment across low resource languages (LRLs). Among them, the Random Forest model performed particularly well, capturing sentiment polarity and handling language-specific nuances effectively. Furthermore, Bidirectional Encoder Representations from Transformers (BERT), a Large Language Model (LLM), is applied to predict context-based sentiment with high accuracy, achieving 99% accuracy and 98% precision, outperforming other models. The BERT predictions were clarified using Explainable AI (XAI), improving transparency and fostering confidence in sentiment classification. Overall, findings demonstrate that the proposed lexicon and machine learning models significantly enhance translation and sentiment analysis for LRLs in South Africa and the DRC, laying a foundation for future AI models that support underrepresented languages, with applications across education, governance, and business in multilingual contexts. This work is part of a PhD proposal in Information Technology at the University of Pretoria, supervised by Dr. Mike Wa Nkongolo and co-supervised by Dr. Phil van Deventer, under the Low-Resource Language Processing Lab in the Department of Informatics

Visit

arxiv.org

Tasks

sentiment analysistext classification

Languages

AfrikaansLuba-KasaiSotho, NorthernZulu

Tags

Computation and LanguageArtificial IntelligenceMachine Learning

Similar

Zero-shot Sentiment Analysis in Low-Resource Languages Using a Multilingual Sentiment LexiconMachine Translation Hallucination Detection for Low and High Resource Languages using Large Language ModelsInteractive Machine Translation with Large Language Models for Low-resource LanguagesUtilizing Multilingual Encoders to Improve Large Language Models for Low-Resource LanguagesDevelopment of a Multilingual Lexicon Based on Sentiment Analysis for Low-Resource LanguagesEvaluating Large Language Models for Low-Resource Multilingual Machine Translation in the Medical Domain

Zero-shot Sentiment Analysis in Low-Resource Languages Using a Multilingual Sentiment Lexicon

Improving multilingual language models capabilities in low-resource languages is generally difficult

Machine Translation Hallucination Detection for Low and High Resource Languages using Large Language Models

Recent advancements in massively multilingual machine translation systems have significantly enhance

Interactive Machine Translation with Large Language Models for Low-resource Languages

Large language models (LLM) have been applied to machine translation with notable success. However,

Utilizing Multilingual Encoders to Improve Large Language Models for Low-Resource Languages

Large Language Models (LLMs) excel in English, but their performance degrades significantly on low-r

Development of a Multilingual Lexicon Based on Sentiment Analysis for Low-Resource Languages

Abstract The multilingual landscape of South Africa and the Democratic Republic of Congo (

Evaluating Large Language Models for Low-Resource Multilingual Machine Translation in the Medical Domain

This dissertation explores neural machine translation (NMT) in multilingual medical domain, with