Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

RideKE: Leveraging Low-Resource, User-Generated Twitter Content for Sentiment and Emotion Detection in Kenyan Code-Switched Dataset

Domain:

natural language processing

Record type:

paperdataset
Creator:
EtoGin
Host:avatar
Social media has become a crucial open-access platform for individuals to express opinions and share experiences. However, leveraging low-resource language data from Twitter is challenging due to scarce, poor-quality content and the major variations in language use, such as slang and code-switching. Identifying tweets in these languages can be difficult as Twitter primarily supports high-resource languages. We analyze Kenyan code-switched data and evaluate four state-of-the-art (SOTA) transformer-based pretrained models for sentiment and emotion classification, using supervised and semi-supervised methods. We detail the methodology behind data collection and annotation, and the challenges encountered during the data curation phase. Our results show that XLM-R outperforms other models; for sentiment analysis, XLM-R supervised model achieves the highest accuracy (69.2\%) and F1 score (66.1\%), XLM-R semi-supervised (67.2\% accuracy, 64.1\% F1 score). In emotion analysis, DistilBERT supervised leads in accuracy (59.8\%) and F1 score (31\%), mBERT semi-supervised (accuracy (59\% and F1 score 26.5\%). AfriBERTa models show the lowest accuracy and F1 scores. All models tend to predict neutral sentiment, with Afri-BERT showing the highest bias and unique sensitivity to empathy emotion. github.com Accepted in WASSA 2024

Visit

arxiv.org

Tasks

code switchingemotion identificationsentiment analysistext classification

Languages

Wasa

Tags

Computation and LanguageArtificial Intelligence

Similar

Identifying Sentiments in Algerian Code-switched User-generated CommentsOffensive Content Detection Via Synthetic Code-Switched TextSAGE: Spliced-Audio Generated Data for Enhancing Foundational Models in Low-Resource Arabic-English Code-Switched Speech RecognitionLeveraging Twitter for Low-Resource Conversational Speech Language ModelingPolitiKweli: A Swahili-English Code-switched Twitter Political Misinformation Classification DatasetSundanese Twitter Dataset for Emotion Classification

Identifying Sentiments in Algerian Code-switched User-generated Comments

We present in this paper our work on Algerian language, an under-resourced North African colloquial Arabic variety, for which we built a comparably large corpus of more than 36,000 code-switched user-generated comments annotated for sentiments. We opted for this da

Offensive Content Detection Via Synthetic Code-Switched Text

The prevalent use of offensive content in social media has become an important reason for concern fo

SAGE: Spliced-Audio Generated Data for Enhancing Foundational Models in Low-Resource Arabic-English Code-Switched Speech Recognition

This paper investigates the performance of various speech SSL models on dialectal Arabic (DA) and Ar

Leveraging Twitter for Low-Resource Conversational Speech Language Modeling

In applications involving conversational speech, data sparsity is a limiting factor in building a be

PolitiKweli: A Swahili-English Code-switched Twitter Political Misinformation Classification Dataset

PolitiKweli: A Swahili-English Code-switched Twitter Political Misinformation Classification Dataset

Poster presented at the Deep Learning Indaba 2023 by CYNTHIA JAYNE AMOL

Sundanese Twitter Dataset for Emotion Classification