Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

A multilingual dataset for offensive language and hate speech detection for hausa, yoruba and igbo languages

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
AliWajMURTALA MUHAMMAD
Hôte:avatar
The proliferation of online offensive language necessitates the development of effective detection mechanisms, especially in multilingual contexts. This study addresses the challenge by developing and introducing novel datasets for offensive language detection in three major Nigerian languages: Hausa, Yoruba, and Igbo. We collected data from Twitter and manually annotated it to create datasets for each of the three languages, using native speakers. We used pre-trained language models to evaluate their efficacy in detecting offensive language in our datasets. The best-performing model achieved an accuracy of 90\%. To further support research in offensive language detection, we plan to make the dataset and our models publicly available. The experimental result was erroneously reported and we also omitted other authors

Visit

arxiv.org

Tasks

hate speech detectiontext classification

Languages

HausaIgboYoruba

Tags

Computation and Language14F05F.2.2; I.2.7

Similaires

Hate Speech and Offensive Language Detection in BengaliAlgerian Dialect Dataset Targeted Hate Speech, Offensive Language and CyberbullyingA deep learning based multilingual hate speech detection for resource scarce languagesCombining FastText and Glove Word Embedding for Offensive and Hate speech Text DetectionAmharic dataset for hate speech detectionAfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages

Hate Speech and Offensive Language Detection in Bengali

Social media often serves as a breeding ground for various hateful and offensive content. Identifyin

Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying

Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying.   * To cite t

A deep learning based multilingual hate speech detection for resource scarce languages

Over the last decade, the increased use of social media has led to an increase in hateful activities

Combining FastText and Glove Word Embedding for Offensive and Hate speech Text Detection

Combining FastText and Glove Word Embedding for Offensive and Hate speech Text Detection

Poster presented at the Deep Learning Indaba 2022 by Nabil BADRI

Amharic dataset for hate speech detection

the dataset is collected from social media such as facebook and telegram. the dataset is further pro

AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages

Hate speech and abusive language are global phenomena that need socio-cultural background knowledge