Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

CyberDTD: A Multimodal Benchmark Dataset for Cyberbullying Detection in Tunisian Dialect

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
BecMekBadEll
Éditeur:
MulجامTunLab
Éditeur:
CCSD
Hôte:avatar
International audience Effective detection of cyberbullying requires understanding both textual and visual signals, including images with embedded text and user generated comments. This need is even more evident in low resource and multilingual environments such as Tunisia. In this context, this paper establishes CyberDTD (Cyberbullying Detection in Tunisian Dialect), a multimodal dataset designed to support research on cyberbullying detection in the Tunisian Dialect (TD). With 10,802 images across five categories, humor, sarcasm, hate, violence, and neutral. We present, to the best of our knowledge, the first cyberbullying dataset in TD. We provide a comprehensive description covering a wide range of online harassment, while also including neutral examples for balanced analysis. Key challenges such as class imbalance, multimodality, and cultural specificity are highlighted. CyberDTD represents an important resource for building and evaluating machine learning models in low-resource settings, supporting the development of more robust and culturally aware cyberbullying detection systems.

Visit

hal.science

Tasks

hate speech detectiontext classification

Languages

Arabic, Tunisian Spoken

Tags

Natural Language Processing (NLP)CyberbullyingTunisian DialectMultimodal Dataset[INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL][INFO.INFO-TI]Computer Science [cs]/Image Processing [eess.IV][INFO.INFO-TT]Computer Science [cs]/Document and Text Processing[STAT.ML]Statistics [stat]/Machine Learning [stat.ML]

Licenses

https://creativecommons.org/licenses/by-nc-nd/4.0/info:eu-repo/semantics/OpenAccess

Similaires

Bully.tn: A Cyberbullying Detection Dataset in Tunisian DialectTunisian dialect cyberbullying detection with ML DL approach and rich feature setTunDC: a public benchmark dataset for sentiment analysis and language modeling in the Tunisian dialectHateTune: Tunisian Dialect Hate Speech Detection DatasetDjaziaH/Cyberbullying-Detection-DataSetDziriFake: A Dataset and Comparative Benchmark of Large Language Models for Fake News Detection In Algerian Dialect

Bully.tn: A Cyberbullying Detection Dataset in Tunisian Dialect

Tunisian dialect cyberbullying detection with ML DL approach and rich feature set

Automatic detection of cyberbullying in the Tunisian dialect is a crucial area of research, particul

TunDC: a public benchmark dataset for sentiment analysis and language modeling in the Tunisian dialect

The development of natural language processing (NLP) applications has increasingly focused on dialec

HateTune: Tunisian Dialect Hate Speech Detection Dataset

DjaziaH/Cyberbullying-Detection-DataSet

The first annotated dataset for the detection of cyberbullying and harassment in the Algerian online

DziriFake: A Dataset and Comparative Benchmark of Large Language Models for Fake News Detection In Algerian Dialect