Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Low-Resource Hate Speech Detection in English-Swahili Code-Switched Text Using Fine-Tuning of Pre-trained Language Models

Domain:

natural language processing

Record type:

datasetpaper
Creator:
Kipkebut, AndrewJep
Publisher:
Int
Host:
The use of social media in East Africa has grown rapidly, and with it, the spread of hate speech has become a serious concern. This problem is even more complex in online spaces where people often switch between English and Swahili within the same sentence or conversation. Such code-switching makes it difficult for existing systems to accurately detect harmful content, especially because there is limited labeled data and much of the language used is informal and context-dependent. This study explores a low-resource approach to detecting hate speech in English and Swahili code-switched text by fine-tuning pre-trained language models. In this work, transformer-based models such as BERT and AfriBERTa are adapted to better understand mixed-language communication. The models are trained on a carefully prepared dataset made up of real social media posts that reflect how people actually write and speak online. These posts are manually labeled to capture both direct and subtle forms of hate speech, including expressions that are influenced by local culture and everyday slang. The findings show that fine-tuned models perform better than traditional machine learning approaches, especially in terms of accuracy and overall detection quality. They are also more effective at handling informal language, abbreviations, and mixed grammar structures. Beyond performance, the study also looks at fairness and bias, emphasizing the need for systems that are sensitive to cultural and linguistic diversity. Overall, this work shows that fine-tuning modern language models can offer a practical and scalable solution for hate speech detection in multilingual environments.

Visit

doi.org

Tasks

code switchinghate speech detectiontext classification

Languages

Swahili

Similar

SWAHILI AND CODE-SWITCHED ENGLISH-SWAHILI POLITICAL HATE SPEECH DETECTION TEXTUAL DATASETFine-Tuning DeepSpeech Speech-To-Text Model for Nigerian English and Yoruba-English Code-Switched SpeechAbusive and Hate speech Classification in Arabic Text Using Pre-trained Language Models and Data AugmentationA Comparative Study of Transformer-based Models for Hate-Speech Detection in English-Kiswahili Code-Switched Social Media TextDisfluent-to-Fluent Tunisian Dialect Speech Translation with Fine-Tuning Pre-trained Language ModelsFine-Tuning Large Language Models with QLoRA for Offensive Language Detection in Roman Urdu-English Code-Mixed Text

SWAHILI AND CODE-SWITCHED ENGLISH-SWAHILI POLITICAL HATE SPEECH DETECTION TEXTUAL DATASET

This dataset consist of swahili and code switched English-Swahili tweets labeled for hate speech typ

Fine-Tuning DeepSpeech Speech-To-Text Model for Nigerian English and Yoruba-English Code-Switched Speech

International audience Speech-to-Text (STT) systems, despite their stellar performanc

Abusive and Hate speech Classification in Arabic Text Using Pre-trained Language Models and Data Augmentation

Hateful content on social media is a worldwide problem that adversely affects not just the targeted

A Comparative Study of Transformer-based Models for Hate-Speech Detection in English-Kiswahili Code-Switched Social Media Text

The transformer architecture, first introduced in 2017 by researchers at Google, has revolutionized

Disfluent-to-Fluent Tunisian Dialect Speech Translation with Fine-Tuning Pre-trained Language Models

Fine-Tuning Large Language Models with QLoRA for Offensive Language Detection in Roman Urdu-English Code-Mixed Text

The use of derogatory terms in languages that employ code mixing, such as Roman Urdu, presents chall