Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Toxicity Detection Dataset in Twi Language

Domaine:

natural language processing

Type de record:

dataset
Créateur:
AliMenAyiAya
Éditeur:
Uni
Éditeur:
Men
Hôte:avatar
This dataset contains 2,001 text entries labeled for toxicity classification. Each entry represents a user-generated comment along with an assigned toxicity label. The dataset is structured into two columns: COMMENT– A text field containing comments written primarily in Akan (Twi). These comments include expressions of gratitude, feedback, conversational messages, and general communication typical of social or online interactions. LABEL– A categorical variable indicating whether the comment is 'toxic' or 'non-toxic'. Current labels present in the dataset: 'non-toxic' (and any others present in the full file, if applicable). Key Features: • Total records: 2,001 • Language: Primarily Akan (Twi) • Classification type: Binary toxicity classification There are no missing values (both columns have 2,001 non-null entries) Data types: ‘COMMENT’: string and ‘LABEL’`: string This dataset can support research in: • Toxic language detection in low-resource languages • Natural Language Processing (NLP) for African languages • Machine learning model training for text classification • Sociolinguistic analysis of online conversational content The File Format is CSV file: Toxicity_dataset.csv It contains two columns: 'COMMENT' and ‘LABEL'

Visit

doi.orgdata.mendeley.com

Tasks

hate speech detectiontext classification

Languages

AkanBwamu, CwiDinka, SoutheasternTwi

Tags

Toxicity

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode