Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages

Domain:

natural language processing

Record type:

paperdataset
Creator:
Muhammad, Shamsuddeen HassanAbdulmumin, IdrisAyeAde
Host:avatar
Hate speech and abusive language are global phenomena that need socio-cultural background knowledge to be understood, identified, and moderated. However, in many regions of the Global South, there have been several documented occurrences of (1) absence of moderation and (2) censorship due to the reliance on keyword spotting out of context. Further, high-profile individuals have frequently been at the center of the moderation process, while large and targeted hate speech campaigns against minorities have been overlooked. These limitations are mainly due to the lack of high-quality data in the local languages and the failure to include local communities in the collection, annotation, and moderation processes. To address this issue, we present AfriHate: a multilingual collection of hate speech and abusive language datasets in 15 African languages. Each instance in AfriHate is annotated by native speakers familiar with the local culture. We report the challenges related to the construction of the datasets and present various classification baseline results with and without using LLMs. The datasets, individual annotations, and hate speech and offensive language lexicons are available on AfriHate: Multilingual Hate…

Visit

arxiv.org

Tasks

hate speech detectiontext classification

Tags

Computation and Language

Similar

AfriHate: Multilingual Hate Speech DatasetIntersectional Bias in Hate Speech and Abusive Language DatasetsRacial Bias in Hate Speech and Abusive Language Detection DatasetsA multilingual dataset for offensive language and hate speech detection for hausa, yoruba and igbo languagesT-HSAB: A Tunisian Hate Speech and Abusive DatasetAfri Code Datasets (A collection of datasets for code generation in African languages)

AfriHate: Multilingual Hate Speech Dataset

Hate speech detection, NLP research Notes / challenges: Limited to 15 African languages

Intersectional Bias in Hate Speech and Abusive Language Datasets

Algorithms are widely applied to detect hate speech and abusive language in social media. We investi

Racial Bias in Hate Speech and Abusive Language Detection Datasets

Technologies for abusive language detection are being developed and applied with little consideratio

A multilingual dataset for offensive language and hate speech detection for hausa, yoruba and igbo languages

The proliferation of online offensive language necessitates the development of effective detection m

T-HSAB: A Tunisian Hate Speech and Abusive Dataset

Afri Code Datasets (A collection of datasets for code generation in African languages)

Training and evaluating Large Language Models (LLMs) for code generation, building AI-powered coding