Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

textdetox/multilingual_toxic_lexicon

Domain:

natural language processing

Record type:

dataset
Creator:
tex
Host:
[2025] The lexicon is extended to new languages! Now also included: Italian, French, Hebrew, Hindi, Japanese, Tatar. The list is used on TextDetox 2025 shared task. [2024] The compilation for 9 languages (English, Russian, Ukrainian, Spanish, German, Amharic, Arabic, Chinese, Hindi) toxic words lists which is used for TextDetox 2024 shared task. The list of original sources: English: link Russian: link Ukrainian: link Spanish: link German: link

Visit

huggingface.co

Tasks

hate speech detectiontext classification

Languages

Amharic

Tags

toxic

Licenses

openrail++

Similar

klamas/multilingual_toxic_lexiconGemDetox at TextDetox CLEF 2025: Enhancing a Massively Multilingual Model for Text Detoxification on Low-resource Languages

klamas/multilingual_toxic_lexicon

[2025] The lexicon is extended to new languages! Now also included: Italian, French, Hebrew, Hindi,

GemDetox at TextDetox CLEF 2025: Enhancing a Massively Multilingual Model for Text Detoxification on Low-resource Languages

As social-media platforms emerge and evolve faster than the regulations meant to oversee them, autom