Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Hate speech Dataset (Somali Language) - Sheet1.csv

Domaine:

natural language processing

Type de record:

dataset
Créateur:
MohMoa
Éditeur:
fig
Hôte:avatar
This dataset contains 8,049 text samples in the Somali language, labeled for hate speech detection. Each entry includes an ID, the original text, and a binary label indicating whether the text contains hate speech (1) or not (0). The dataset is the first publicly available resource of its kind for Somali, a low-resource language in NLP, and is intended to support research in natural language processing, computational linguistics, and machine learning for hate speech detection and content moderation.Structure:ID: Unique identifier for each text sample.Text: The Somali-language text content.Label: Binary classification label — 1 for hate speech, 0 for non-hate speech.Modification and Unnamed: Auxiliary columns with partial or empty entries (may be ignored for modeling).Potential Uses:Training and evaluating hate speech detection models.Research on low-resource language processing.Sociolinguistic analysis of harmful language in Somali.Citation: If you use this dataset, please cite it using the DOI provided on this Figshare record.License: (CC BY 4.0 for open use with attribution)Acknowledgment: This dataset was created to advance open research in Somali language technologies and to address gaps in low-resource NLP.

Visit

doi.orgfigshare.com

Tasks

hate speech detectiontext classification

Languages

Somali

Tags

Natural language processing

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode