Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

franciellevargas/HausaHate

Domaine:

natural language processing

Type de record:

dataset
Créateur:
fra
Hôte:
HausaHate is a benchmark dataset for Hausa hate speech detection task. it was extracted from West African Facebook pages and comprises 2,000 comments annotated according to a binary class (offensive and non-offensive) and hate speech targets (race, gender and none). HausaHate: A Benchmark Dataset for Hausa Hate Speech Detection In African countries, the hate speech phenomenon is especially serious due to a historical problem regarding ethnic conflicts. Specifically, the Western region still lacks more research on hate speech focusing on its indigenous languages. Moreover, as most of the existing hate speech data resources are developed for the English language, the research and development of hate speech technologies for African indigenous languages are less developed. To fill this relevant gap, we introduce the first expert annotated corpus of Facebook comments for Hausa hate speech detection. The corpus titled HausaHate comprises 2,000 comments extracted from Western African Facebook pages and manually annotated by three Hausa native speakers, who are also NLP experts. Our corpus was annotated using two different layers. We first labeled each comment according to a binary classification: offensive versus non-offensive. Then, offensive comments were also labeled according to hate speech targets: race, gender and none. Lastly, a baseline model using fine-tuned LLM for Hausa hate speech detection is presented, highlighting the challenges of hate speech detection tasks for indigenous languages in Africa, as well as future advances. The following table describes in detail the HausaHate categories and documents: | Offensive| Non-Offensive | Total Comments | | :--- | :---: | ---: | | 678 | 1,322 | 2,000 | | Race | Gender | Non-Target | Total | | :--- | :---: | ---: | ---: | | 391 | 65 | 222 | 678 | What the following is the list of collaborators and authors this project: ### Lead Authors - **Francielle Vargas** — University of São Paulo, Brazil - **Shamsuddeen H. Muhammad** — Imperial College London, UK ### Researchers - **Ibrahim Said Ahmad** — Northeastern University, USA - **Diego Alves** — Saarland University, Germany - **Idris Abdulmumin** — University of Pre …

Visit

github.com

Tasks

hate speech detectiontext classification

Languages

Hausa

Tags

benchmarkcorpusdatasethate-speechhausa-nlplow-resource-languagesmachine-learningnatural-language-processingnlp-machine-learningoffensive-language

Similaires

franciellevargas/HausaHatefranciellevargas/HausaHate: v3.0.0

franciellevargas/HausaHate

HausaHate is the first Hausa hate speech dataset extracted from West African Facebook pages.

franciellevargas/HausaHate: v3.0.0

HausaHate is a benchmark dataset for Hausa hate speech detection task. it was extracted from West Af