Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Balanced Shona Hate Speech Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
oma
Host:
This dataset contains 2,000 balanced examples of Shona text classified into four categories: NEUTRAL, OFFENSIVE, CONTEXTUAL, and HATE. Label Source Count NEUTRAL Literary novel (Imbwa Yemunhu by Ignatius T. Mabasa) 500 OFFENSIVE Synthetic template-based generation 500 CONTEXTUAL Synthetic (quoted hate speech, not endorsed) 500 HATE

Visit

huggingface.co

Tasks

hate speech detectiontext classification

Languages

Shona

Tags

shonahate-speechoffensive-languageafrican-nlplow-resource

Licenses

apache-2.0

Similar

omanyasa/shona-hate-speech-modelAmharic Hate Speech DatasetKiswahili Hate speech datasetshunyalabs/shona-speech-datasetShona Speech Dataset (SNA)Egyptian Arabic Hate Speech Dataset

omanyasa/shona-hate-speech-model

Amharic Hate Speech Dataset

Introduction

Kiswahili Hate speech dataset

Kiswahili Hate Speech Dataset (KHS-2026) Overview KHS-2026 is a monolingual Kiswahili dataset create

shunyalabs/shona-speech-dataset

Shona Speech Dataset (SNA)

A cleaned, metadata-rich Shona (sna) speech dataset prepared through a reproducible data engineering

Egyptian Arabic Hate Speech Dataset

Author: IbrahimAmin, Mostafa Abbas, Rany Hatem, Andrew Ihab, Mohamed Waleed Fahkr License: MIT Paper