Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Resources Building for Arabic Harmful Online Content: Survey

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
ChaZitDjoRah
Éditeur:
UniLabTeccen
Éditeur:
CCSDAtl
Hôte:avatar
International audience

Users of social networks and Internet sites face numerous challenges. Problems such as fake news, satire, rumors, misinformation, misleading information, cyberbullying, spam content, offensive language, hate, and offensive speech fall under the category of harmful online content (HOC). This danger has taken advantage of social media's popularity and the abundance of news that spreads quickly, causing problems for individuals and society. Moreover, to combat this danger, researchers in the AI domain have persistently advanced and proposed novel approaches across various domains. Given the progress made in this work, choosing data to evaluate their approaches was always a challenge. Our contribution aims to identify the process and criteria for creating a high-quality dataset for HOC detection, primarily in the Arabic news domain. Therefore, we have collected a list of existing and available Arabic datasets, identified their characteristics, and determined the purpose of their creation. Researchers can use our study's results as a reference to choose an appropriate dataset for their future research.

Visit

hal.science

Tasks

hate speech detectiontext classification

Tags

Harmful online contentFake newsOffensive languageNLPArabic datasets[INFO]Computer Science [cs]

Licenses

https://creativecommons.org/licenses/by-nc/4.0/info:eu-repo/semantics/OpenAccess