Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A comparative study for classical machine learning models for swahili social media sentiment analysis

Domain:

natural language processing

Record type:

paper
Creator:
MahDav
Publisher:
Uni
Host:
Despite sentiment analysis being one of the most popular applications in Natural Language Processing (NLP), most studies are skewed towards languages with a rich corpus (language database). Less emphasis has been placed on low-resource languages like Swahili. Swahili is the official language of the African Union and of 4 countries in East Africa, and is spoken by many people on the African continent. This study performed sentiment analysis using 3,000 tweets hosted on the Zindi Africa platform. Data was processed using a term frequency-inverse document frequency vectorization method, and five classical machine learning algorithms (RandomForest, XgBoost, and CatBoost, HistogramGradientBoost, LightGradientBoos) were trained and evaluated using the collected tweets. We found that CatBoost produced the highest performance in general compared to other classical models, with 0.610 accuracy, 0.470 F1 score, 0.522 Precision and 0.462 Recall. The F1-score of 0.47 indicates modest performance and reflects the challenges posed by the small dataset and the complexity of Swahili sentiment analysis. This study offers a comprehensive overview of the relative performance of various classical machine learning models applied to Swahili social media sentiment data. These insights can help researchers make informed choices when selecting appropriate classical machine learning algorithms for sentiment analysis in a similar context.

Visit

doi.org

Tasks

sentiment analysistext classification

Languages

Swahili

Licenses

http://creativecommons.org/licenses/by/4.0/

Similar

Comparative Analysis of Machine Learning Algorithms for Sentiment Analysis of Multilingual Nigerian Social Media CommentsComparative Analysis of Machine Learning and Deep Learning Models for Sentiment Analysis in Somali LanguageMosesKKhoza/Swahili-Social-Media-Sentiment-AnalysisIntroducing a Swahili social media sentiment analysis dataset for the telecom industryComparative Analysis of Machine Learning Models for Detecting Mobile Messaging Spam In Swahili SMSQualitative Study on Social Media Sentiment Analysis for Evidence-Based Policy-Making in Electronic Voting in Nigeria: A Machine Learning Approach

Comparative Analysis of Machine Learning Algorithms for Sentiment Analysis of Multilingual Nigerian Social Media Comments

This study probes into sentiment analysis within the multilingual landscape of Nigerian social media

Comparative Analysis of Machine Learning and Deep Learning Models for Sentiment Analysis in Somali Language

MosesKKhoza/Swahili-Social-Media-Sentiment-Analysis

## Swahili Sentiment Analysis Classifying Swahili tweets into positive, negative, and neutral sentim

Introducing a Swahili social media sentiment analysis dataset for the telecom industry

Swahili is the most widely spoken language in Africa with over 200 million speakers. Despite its popularity in the continent, there is insufficient NLP research conducted on the language. The shortage of high-quality annotated datasets is attributed to this. In t

Comparative Analysis of Machine Learning Models for Detecting Mobile Messaging Spam In Swahili SMS

Qualitative Study on Social Media Sentiment Analysis for Evidence-Based Policy-Making in Electronic Voting in Nigeria: A Machine Learning Approach

Electronic voting (e-voting) has the potential to improve voter access, reduce logistical issues, an