Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

TriLex: A Framework for Multilingual Sentiment Analysis in Low-Resource South African Languages

Domain:

natural language processing

Record type:

paperdataset
Creator:
NkoVorWarNai
Host:avatar
Low-resource African languages remain underrepresented in sentiment analysis, limiting both lexical coverage and the performance of multilingual Natural Language Processing (NLP) systems. This study proposes TriLex, a three-stage retrieval augmented framework that unifies corpus-based extraction, cross lingual mapping, and retrieval augmented generation (RAG) driven lexical refinement to systematically expand sentiment lexicons for low-resource languages. Using the enriched lexicon, the performance of two prominent African pretrained language models (AfroXLMR and AfriBERTa) is evaluated across multiple case studies. Results demonstrate that AfroXLMR delivers superior performance, achieving F1-scores above 80% for isiXhosa and isiZulu and exhibiting strong cross-lingual stability. Although AfriBERTa lacks pre-training on these target languages, it still achieves reliable F1-scores around 64%, validating its utility in computationally constrained settings. Both models outperform traditional machine learning baselines, and ensemble analyses further enhance precision and robustness. The findings establish TriLex as a scalable and effective framework for multilingual sentiment lexicon expansion and sentiment modeling in low-resource South African languages.

Visit

arxiv.org

Tasks

sentiment analysistext classification

Languages

XhosaZulu

Tags

Computation and Language

Similar

Zero-shot Sentiment Analysis in Low-Resource Languages Using a Multilingual Sentiment LexiconHYBRID DEEP LEARNING MODELS FOR MULTILINGUAL SENTIMENT ANALYSIS IN LOW-RESOURCE LANGUAGESDevelopment of a Multilingual Lexicon Based on Sentiment Analysis for Low-Resource LanguagesRonaldKato/Luganda-Language-Analysis-Framework-for-Low-Resource-African-LanguagesAfrican Voices: Multilingual Speech Dataset for Low-Resource African LanguagesEnhancing Sentiment Analysis in Amharic: Leveraging Transformer-Based Language Model for Low-Resource African Languages

Zero-shot Sentiment Analysis in Low-Resource Languages Using a Multilingual Sentiment Lexicon

Improving multilingual language models capabilities in low-resource languages is generally difficult

HYBRID DEEP LEARNING MODELS FOR MULTILINGUAL SENTIMENT ANALYSIS IN LOW-RESOURCE LANGUAGES

Multilingual sentiment analysis poses significant challenges, especially in the context of languages

Development of a Multilingual Lexicon Based on Sentiment Analysis for Low-Resource Languages

Abstract The multilingual landscape of South Africa and the Democratic Republic of Congo (

RonaldKato/Luganda-Language-Analysis-Framework-for-Low-Resource-African-Languages

This repository implements a complete NLP pipeline for analyzing Luganda, a Bantu language spoken by

African Voices: Multilingual Speech Dataset for Low-Resource African Languages

A large-scale multilingual speech dataset developed by Data Science Nigeria. Contains more than 3,000 hours of transcribed audio across four Nigerian languages: Hausa, Igbo, Nigerian Pidgin, and Yorùbá. The dataset supports Automatic Speech Recognition (ASR) and sp

Enhancing Sentiment Analysis in Amharic: Leveraging Transformer-Based Language Model for Low-Resource African Languages