Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A Comparative Study of Embedding Models for Offensive Language Detection in Nigeria

Domain:

natural language processing

Record type:

paperdataset
Creator:
SamSha
Publisher:
Dep
Host:
Offensive language detection has become a pivotal challenge in Natural Language Processing (NLP), particularly in the context of low-resource languages. Existing models, predominantly trained on high-resource languages such as English, often struggle to generalize across linguistically and culturally diverse contexts. This study addresses the limitation by developing and introducing novel datasets for the detection of offensive and hate speech in three major Nigerian languages, Hausa, Yoruba, and Igbo. Data was collected from Twitter and manually annotated using native speakers. The study leverages Afro-XLM-R embedding pre-trained on African language corpora combined with a Convolutional Neural Network (CNN) architecture for robust classification. To evaluate its effectiveness, the proposed hybrid model was benchmarked against three alternative CNN-based models utilizing FastText, word2Vec, and mBERT embedding. Experimental results demonstrate the superiority of the Afro-XLM-R + CNN approach in capturing nuanced linguistic features across the selected languages, highlighting its potential for scalable offensive language detection in low-resource settings.

Visit

doi.org

Tasks

embeddingshate speech detectiontext classification

Languages

HausaYoruba

Similar

Adversarial Evaluation of Large Language Models for Building Robust Offensive Language Detection in Moroccan ArabicComparative Performance of Ensemble Machine Learning for Arabic Cyberbullying and Offensive Language DetectionOffensive Language Detection in ArabiziCombining FastText and Glove Word Embedding for Offensive and Hate speech Text DetectionTRANSFORMER-BASED SPAM DETECTION FOR UZBEK LANGUAGE: A COMPARATIVE STUDY WITH TRADITIONAL MACHINE LEARNING MODELSrayenFathallah/Tunisian-Offensive-Language-Detection

Adversarial Evaluation of Large Language Models for Building Robust Offensive Language Detection in Moroccan Arabic

Offensive language detection is crucial for ensuring safe and inclusive digital environments. Identi

Comparative Performance of Ensemble Machine Learning for Arabic Cyberbullying and Offensive Language Detection

Abstract In recent years, research on abusive language and cyberbullying detection have ga

Offensive Language Detection in Arabizi

Detecting offensive language in under-resourced languages presents a significant real-world challeng

Combining FastText and Glove Word Embedding for Offensive and Hate speech Text Detection

Combining FastText and Glove Word Embedding for Offensive and Hate speech Text Detection

Poster presented at the Deep Learning Indaba 2022 by Nabil BADRI

TRANSFORMER-BASED SPAM DETECTION FOR UZBEK LANGUAGE: A COMPARATIVE STUDY WITH TRADITIONAL MACHINE LEARNING MODELS

Spam detection remains a critical challenge in natural language processing, particularly for low-res

rayenFathallah/Tunisian-Offensive-Language-Detection

This repository contains the implementation of a real-time content filtering system designed to dete