Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A Natural Language Processing Framework for Toxicity Detection in Low-Resource Languages: A Case Study on the Twi Language

Domain:

natural language processing

Record type:

dataset
Creator:
AliMenAyiAya
Editor:
Uni
Publisher:
Men
Host:avatar
This dataset contains 2,001 text entries labeled for toxicity classification. Each entry represents a user-generated comment along with an assigned toxicity label. The dataset is structured into two columns: COMMENT– A text field containing comments written primarily in Akan (Twi). These comments include expressions of gratitude, feedback, conversational messages, and general communication typical of social or online interactions. LABEL– A categorical variable indicating whether the comment is 'toxic' or 'non-toxic'. Current labels present in the dataset: 'non-toxic' (and any others present in the full file, if applicable). Key Features: • Total records: 2,001 • Language: Primarily Akan (Twi) • Classification type: Binary toxicity classification There are no missing values (both columns have 2,001 non-null entries) Data types: ‘COMMENT’: string and ‘LABEL’`: string This dataset can support research in: • Toxic language detection in low-resource languages • Natural Language Processing (NLP) for African languages • Machine learning model training for text classification • Sociolinguistic analysis of online conversational content The File Format is CSV file: Toxicity_dataset.csv It contains two columns: 'COMMENT' and ‘LABEL'

Visit

doi.orgdata.mendeley.com

Tasks

hate speech detectiontext classification

Languages

AkanBwamu, CwiDinka, SoutheasternTwi

Tags

Toxicity

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Natural Language Processing in Low-Resource Languages: Progress and ProspectsToxicity Detection Dataset in Twi LanguageNatural Language Processing (NLP) tools - multilingual and low-resource languagesA neural approach for inducing multilingual resources and natural language processing tools for low-resource languagesChallenges and Opportunities in Natural Language Processing for African Languages in Morocco: A Methodological FrameworkNatural language processing for African languages

Natural Language Processing in Low-Resource Languages: Progress and Prospects

Low-resource languageslanguages with limited annotated corpora, lexicons, and digital resourcespose

Toxicity Detection Dataset in Twi Language

This dataset contains 2,001 text entries labeled for toxicity classification. Each entry represents

Natural Language Processing (NLP) tools - multilingual and low-resource languages

A neural approach for inducing multilingual resources and natural language processing tools for low-resource languages

Abstract This work focuses on the rapid development of linguistic annotation tools for low-resource

Challenges and Opportunities in Natural Language Processing for African Languages in Morocco: A Methodological Framework

Natural Language Processing (NLP) has seen significant advancements in processing languages

Natural language processing for African languages

Recent advances in pre-training of word embeddings and language models leverage large amounts of unlabelled texts and self-supervised learning to learn distributed representations that have significantly improved the performance of deep learning models on a large v