Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Text Detoxification in isiXhosa and Yorùbá Dataset

Domain:

natural language processing

Record type:

datasetpaper
Creator:
Agb
Publisher:
Men
Host:avatar
A parallel dataset of toxic and detoxified sentence pairs (178 pairs each) was manually generated for isiXhosa and Yorùbá, which included a wide range of linguistic and communicative forms such as direct insults, implicit hostility, sarcasm, emotional outbursts, and culturally specific slurs.

Visit

data.mendeley.com

Languages

XhosaYoruba

Tags

Computational LinguisticsNatural Language ProcessingMachine LearningSubsaharan AfricaSouth AfricaNigeria

Licenses

Creative Commons Attribution 4.0 Internationalhttp://creativecommons.org/licenses/by/4.0