Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

AraTox: A Multi-Dialect, Multi-Label Arabic Dataset for Toxicity Detection

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Als
Éditeur:
Men
Hôte:avatar
AraTox is a multi-dialect, multi-label Arabic dataset for toxicity detection. It contains annotated Arabic text representing multiple varieties, including Gulf, Levantine, Nile Basin, North African, Yemeni, and Modern Standard Arabic (MSA). The dataset is designed to support research in toxic language detection, multi-label classification, and the benchmarking of Arabic NLP models across dialects. This dataset accompanies the following publication: Aratox: A multi-dialect, multi-label Arabic dataset and model benchmark for toxicity detection If you use this dataset, please cite: Alshargi, F., Abulohoom, A., Yagi, S., Jabr, F., Lulu, L., & Elnagar, A. (2026). Aratox: A multi-dialect, multi-label Arabic dataset and model benchmark for toxicity detection. Language Resources & Evaluation, 60, 39. doi.org

Visit

doi.orgdata.mendeley.com

Tasks

hate speech detectiontext classification

Tags

Computer ScienceNatural Language ProcessingArabic LanguageArtificial Intelligence Model

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode