Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Language-Aware Cyberbullying Detection in Low-Resource Settings Using Swahili Social Media Data

Domaine:

natural language processing

Type de record:

datasetpaper
Créateur:
KarAdaEvaJum
Éditeur:
Eas
Hôte:
Cyberbullying on social media has emerged as a growing concern affecting online safety, particularly among Swahili-speaking communities that lack robust automatic detection tools. This paper aims to evaluate the effectiveness of machine learning models in detecting cyberbullying in Swahili social media content. The study’s key contributions include the creation of a Swahili text corpus sourced from social media, pre-processing the dataset to ensure linguistic and structural consistency, training natural language processing models on the labelled data and evaluating model performance using accuracy, precision, recall, and F1-score. Four supervised machine learning algorithms, such as Support Vector Machine (SVM), Naïve Bayes (NB), Decision Tree (DT), and Logistic Regression (LR), were implemented and tested with two feature extraction methods: Term Frequency–Inverse Document Frequency (TF-IDF) and Count Vectorizer (CV). Experimental results show that the Decision Tree model with TF-IDF achieved the highest accuracy of 96.40%, precision of 90.58%, recall of 87.49%, and F1-score of 89.01%, outperforming all other models. These findings demonstrate the feasibility of developing efficient and language-specific cyberbullying detection systems for Swahili. Future work recommends expanding the dataset, incorporating deep learning and transformer-based models and developing culturally aware Swahili lexicons to enhance accuracy and contextual understanding. The study contributes to advancing multilingual natural language processing and promoting safer digital interactions in low-resource linguistic environments.

Visit

doi.org

Tasks

hate speech detectiontext classification

Languages

Swahili

Licenses

http://creativecommons.org/licenses/by/4.0

Similaires

Optimizing Uncertainty-Aware Deep Learning for On-the-Edge Murmur Detection in Low-Resource Settingsglorybagai/Domain-Shift-Aware-Parameter-Efficient-Fine-Tuning-for-Malaria-Detection-in-Low-Resource-SettingsUsing Machine Learning for Medical Error Detection in Low-Resource SettingsSwahili Social Media DataA Multi-Task Benchmark for Abusive Language Detection in Low-Resource SettingsMemeGuard: Transformer-Based Fusion for Multimodal Propaganda Detection in Low-Resource Social Media Memes

Optimizing Uncertainty-Aware Deep Learning for On-the-Edge Murmur Detection in Low-Resource Settings

Early and reliable detection of heart murmurs is essential for the timely diagnosis of cardiovascula

glorybagai/Domain-Shift-Aware-Parameter-Efficient-Fine-Tuning-for-Malaria-Detection-in-Low-Resource-Settings

# Domain-Shift-Aware-Parameter-Efficient-Fine-Tuning-for-Malaria-Detection-in-Low-Resource-Settings

Using Machine Learning for Medical Error Detection in Low-Resource Settings

Abstract Medication errors during surgical procedures pose significant risks to pa

Swahili Social Media Data

A dataset on customer feedback and corresponding sentiment scores across various telecommunications companies in Tanzania, as shared on social media platforms.

A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings

Content moderation research has recently made significant advances, but remains limited in serving t

MemeGuard: Transformer-Based Fusion for Multimodal Propaganda Detection in Low-Resource Social Media Memes

Memes are now a common means of communication on social media. Their humor and short format help mes