Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

A Comparative Study of Transformer-based Models for Hate-Speech Detection in English-Kiswahili Code-Switched Social Media Text

Domaine:

natural language processing

Type de record:

paper
Éditeur:
The
Hôte:
The transformer architecture, first introduced in 2017 by researchers at Google, has revolutionized natural language processing in various tasks, including text classification. This architecture formed the basis of future models such as those used in hate speech detection in code-switched text. In this research, we conduct a comparative study of transformer-based models for hate speech detection in English-Kiswahili code-switched text. First, the models were compared as feature extractors using a traditional classifier and then as end-to-end classifiers. The three multilingual transformer-based models compared include mBERT, mDistilBERT and XLM-RoBERTa, using SVM as the traditional classifier for the extracted features. The HateSpeech_Kenya dataset, sourced from Kaggle, was utilized in this study. As a feature extractor, mBERT’s hidden states trained the highest-performing SVM with an accuracy of 0.5461 and a macro f1 score of 0.40. Among the three models evaluated, XLM-RoBERTa achieved the highest accuracy of 0.6069 and a macro f1 score of 0.49 on a balanced dataset. In contrast, mBERT achieved the highest accuracy of 0.7820 and a macro f1 score of 0.53 on an imbalanced dataset. The comparative study establishes that using transformer-based models as end-to-end classifiers generally performs better than using them as feature extractors with traditional classifiers. This is because directly training the models allows them to learn more task-specific features. Furthermore, the varying performance across balanced and imbalanced datasets highlights the need for careful model selection based on the dataset characteristics and specific task requirements.

Visit

doi.org

Tasks

code switchinghate speech detectiontext classification

Languages

SwahiliSwahili, CoastalSwahili, Congo

Similaires

A Context-Aware and Target-Adaptive Multilingual Framework for Hate Speech Detection in Code-Switched Social Media TextTransformer-based Text Generation for Code-Switched Sepedi-English NewsLow-Resource Hate Speech Detection in English-Swahili Code-Switched Text Using Fine-Tuning of Pre-trained Language ModelsPsychosocial Features for Hate Speech Detection in Code-switched TextsSWAHILI AND CODE-SWITCHED ENGLISH-SWAHILI POLITICAL HATE SPEECH DETECTION TEXTUAL DATASETHATE SPEECH DETECTION ON SOCIAL MEDIA FOR AMHARIC TEXT USING DEEP LEARNING APPROACH

A Context-Aware and Target-Adaptive Multilingual Framework for Hate Speech Detection in Code-Switched Social Media Text

The rapid expansion of social media has accelerated the spread of hate speech, particularly within m

Transformer-based Text Generation for Code-Switched Sepedi-English News

Code-switched data is rarely available in written form and this makes the development of la

Low-Resource Hate Speech Detection in English-Swahili Code-Switched Text Using Fine-Tuning of Pre-trained Language Models

The use of social media in East Africa has grown rapidly, and with it, the spread of hate speech has

Psychosocial Features for Hate Speech Detection in Code-switched Texts

This study examines the problem of hate speech identification in codeswitched text from social media

SWAHILI AND CODE-SWITCHED ENGLISH-SWAHILI POLITICAL HATE SPEECH DETECTION TEXTUAL DATASET

This dataset consist of swahili and code switched English-Swahili tweets labeled for hate speech typ

HATE SPEECH DETECTION ON SOCIAL MEDIA FOR AMHARIC TEXT USING DEEP LEARNING APPROACH

HATE SPEECH DETECTION ON SOCIAL MEDIA FOR AMHARIC TEXT USING DEEP LEARNING APPROACH