Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

MuhammadAyob/Multilingual-Sentiment-Analysis-Lexicon-Expansion

Domain:

natural language processing

Record type:

project
Creator:
Muh
Host:
This study expands a bilingual lexicon to include English, Afrikaans, Zulu, Xhosa, and Sesotho for multilingual sentiment analysis in South African languages. SMOTE preprocessing balanced sentiment classes, and Random Forest achieved the best performance. Ensemble models show promise for low-resource languages, enabling culturally aware tools. # Multilingual Sentiment Analysis and Lexicon Expansion for South African Languages This project explores the development of a multilingual sentiment analysis tool for South African languages. It focuses on expanding an existing bilingual lexicon (French-Ciluba) to include English, Afrikaans, Zulu, Xhosa, and Sesotho and evaluates the effectiveness of various machine learning models for sentiment classification. ## Overview South Africa's linguistic diversity presents unique challenges for sentiment analysis. This study addresses the underrepresentation of indigenous languages in NLP by: - Expanding a lexicon to support sentiment classification in multiple languages. - Applying machine learning models to classify sentiments across languages. - Testing translation fidelity to maintain accurate sentiment interpretation. ### Key Features: - **Languages Covered:** English, Afrikaans, Zulu, Xhosa, Sesotho, French, and Ciluba. - **Machine Learning Models:** Random Forest, SVM, Logistic Regression, Decision Tree, and Naive Bayes. - **Preprocessing Techniques:** SMOTE for class balancing, sentiment normalization, and translation validation (manual review and back-translation). - **Evaluation Metrics:** Accuracy, F1 score, precision, recall, and ROC-AUC scores. ## Highlights of Results - **Best Performing Model:** Random Forest achieved the highest accuracy and F1 scores, effectively capturing the complexity of multilingual data. - **Translation Consistency:** Sentiment polarity and intensity were largely preserved across translations, with minor variations due to cultural and linguistic nuances. - **Challenges:** Class imbalance, limited sentiment-rich adjectives in the lexicon, and difficulty in handling nuanced cultural expressions. ## Methodology 1. **Data Preparation:** - Cleaned and normalized text data. - Balanced sentiment classes using SMOTE. - Emphasized sentiment-rich parts of speech (e.g., adjectives). 2. **Model Training and Evaluation:** - Models were …

Visit

github.com

Tasks

sentiment analysistext classification

Languages

AfrikaansLuba-KasaiSotho, SouthernXhosaZulu