Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Introducing a Swahili social media sentiment analysis dataset for the telecom industry

Domain:

natural language processing

Record type:

paper

Swahili is the most widely spoken language in Africa with over 200 million speakers. Despite its popularity in the continent, there is insufficient NLP research conducted on the language. The shortage of high-quality annotated datasets is attributed to this.

In this paper, we introduce a Swahili dataset collected…. The dataset comprised of a comprehensive collection of 8.7K tweets on products and services offered by telecommunications companies based in Tanzania. The tweets on the dataset are annotated manually by Swahili native speakers into three sentiments (Positive, Negative and Neutral). We have provided a detailed description of the steps involved in gathering and annotating the tweets, encompassing an elaborate account of the data collection method, annotation process, and dataset statistics.

We tested the suitability of the developed dataset using five sentiment-classical machine learning models producing F1-scores ranging from 0.6889 to 0.7522 and 5 pre-trained transformer models producing F1-scores ranging from 0.7001 to 0.7306. Further, within this scholarly research paper, we expound upon the challenges encountered during the data collection and annotation processes. These challenges encompass bilingual tweets, the translation of emojis, the absence of Swahili language recognition by the Twitter platform, as well as the intricacies arising from Swahili words or phrases with multiple contextual meanings and informal vocabulary slang, and hashtag misclassification.

Visit

link.springer.comAccess on ResearchGate

Connected records

dataset

Tasks

sentiment analysistext classification

Languages

Swahili

Tags

telecomsocial mediasentimenttanzania

Similar

MosesKKhoza/Swahili-Social-Media-Sentiment-AnalysisIntroducing A large Tunisian Arabizi Dialectal Dataset for Sentiment AnalysisA comparative study for classical machine learning models for swahili social media sentiment analysiselectricsheepafrica/africa-synth-telecom-social-media-sentiment-datasets-nigeriaEnhanced BERT for tourism sentiment: A social media monitoring system for Ugandan tourism industrymahmoudsegni/Social-Media-Sentiment-Analysis-for-Tunisian-Arabizi

MosesKKhoza/Swahili-Social-Media-Sentiment-Analysis

## Swahili Sentiment Analysis Classifying Swahili tweets into positive, negative, and neutral sentim

Introducing A large Tunisian Arabizi Dialectal Dataset for Sentiment Analysis

On various Social Media platforms, people tend to use the informal way to communicate, or write posts and comments: their local dialects. In Africa, more than 1500 dialects and languages exist. Particularly, Tunisians talk and write informally using Latin letters a

A comparative study for classical machine learning models for swahili social media sentiment analysis

Despite sentiment analysis being one of the most popular applications in Natural Language Processing

electricsheepafrica/africa-synth-telecom-social-media-sentiment-datasets-nigeria

Customer posts and complaints from social media platforms Category: Customer Experience and Sentime

Enhanced BERT for tourism sentiment: A social media monitoring system for Ugandan tourism industry

The Enhanced BERT for Tourism Sentiment (EBTS) serves as an advanced sentiment analysis module which

mahmoudsegni/Social-Media-Sentiment-Analysis-for-Tunisian-Arabizi

On social media, Arabic speakers tend to express themselves in their own local dialect. To do so, Tu