Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Modelling Misinformation in Swahili-English Code-switched Texts

Domain:

natural language processing

Record type:

datasetmodel
Creator:
MasCynLilJam
Publisher:
MEC
Host:
Code-switching, which is the mixing of words or phrases from multiple, grammatically distinct languages, introduces semantic and syntactic complexities to sentences which complicate automated text classification. Despite code-switching being a common occurrence in informal text-based communication among most bilingual or multilingual users of digital spaces, its use to spread misinformation is relatively less explored. In Kenya, for instance, the use of code-switched Swahili-English is prevalent on social media. Our main objective in this paper was to systematically re- view code-switching, particularly the use of Swahili-English code-switching to spread misinformation on social media in the Kenyan context. Additionally, we aimed at pre-processing a Swahili-English code-switched dataset and developing a misinformation classification model trained on this dataset. We discuss the process we took to develop the code- switched Swahili-English misinformation classification model. The model was trained and tested using the PolitiKweli dataset which is the first Swahili-English code-switched dataset curated for misinformation classification. The dataset was collected from Twitter (now X) social media platform, focusing on text posted during the electioneering period of the 2022 general elections in Kenya. The study experimented with two types of word embeddings - GloVe and FastText. FastText uses character n-gram representations that help generate meaningful vectors for rare and unseen words in the code-switched dataset. We experimented with both the classical machine learning algorithms and deep learning algo- rithms. Bidirectional Long Short-Term Memory Networks (BiLSTM) algorithm showed the best performance with an f-score of 0.89. The model was able to classify code-switched Swahili-English political misinformation text as fake, fact or neutral. This study contributes to recent research efforts in developing language models for low-resource languages.

Visit

doi.org

Tasks

code switchingtext classification

Languages

Swahili

Similar

PolitiKweli: A Swahili-English Code-switched Twitter Political Misinformation Classification DatasetSWAHILI AND CODE-SWITCHED ENGLISH-SWAHILI POLITICAL HATE SPEECH DETECTION TEXTUAL DATASETCode-switched English Pronunciation Modeling for Swahili Spoken Term DetectionHausa-English Code-Switched DatasetDetecting Propaganda Techniques in Code-Switched Social Media TextsChichewa-English Code-Switched Speech Dataset

PolitiKweli: A Swahili-English Code-switched Twitter Political Misinformation Classification Dataset

PolitiKweli: A Swahili-English Code-switched Twitter Political Misinformation Classification Dataset

Poster presented at the Deep Learning Indaba 2023 by CYNTHIA JAYNE AMOL

SWAHILI AND CODE-SWITCHED ENGLISH-SWAHILI POLITICAL HATE SPEECH DETECTION TEXTUAL DATASET

This dataset consist of swahili and code switched English-Swahili tweets labeled for hate speech typ

Code-switched English Pronunciation Modeling for Swahili Spoken Term Detection

Hausa-English Code-Switched Dataset

Overview The Hausa-English Code-Switched Dataset contains comments collected from Facebook, Instagram, YouTube, and Twitter. These comments exhibit code-switching between Hausa and English, providing a rich resource for linguistic research, natural language proces

Detecting Propaganda Techniques in Code-Switched Social Media Texts

Propaganda is a planned persuasive form of communication whose goal is to influence the opinions and

Chichewa-English Code-Switched Speech Dataset

A speech dataset containing 247 audio recordings of Chichewa-English code-switched phrases. Code-swi