Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Adversarial Evaluation of Large Language Models for Building Robust Offensive Language Detection in Moroccan Arabic

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
SouKanAsmEl
Éditeur:
MDP
Hôte:
Offensive language detection is crucial for ensuring safe and inclusive digital environments. Identifying harmful content protects users and supports healthier online interactions. Despite advances in transformer-based models, particularly Large Language Models (LLMs), their application to this task remains underexplored for low-resource languages such as Moroccan Arabic, especially compared with high-resource languages. This study evaluates the performance of various open- and closed-source LLMs for offensive language detection in Moroccan Darija. The evaluated models include general-purpose LLMs such as LLaMA, Mistral, and Gemma, as well as Arabic-focused models such as ArabianGPT, Falcon Arabic, and Atlas-Chat. We also experiment with reasoning models such as DeepSeek and GPT-4. Beyond traditional evaluation metrics, we investigate the robustness of these LLMs and examine the impact of adversarial training on their performance. Moreover, we contribute to the field by creating a large, high-quality dataset. Our evaluation revealed that GPT-4o Mini achieved the best overall performance, reaching an F1-score of 88%. However, robustness testing under black-box and white-box adversarial attacks exposed notable vulnerabilities, with attack success rates reaching 30%, thereby highlighting the need for enhancement. Despite the complex morphology and linguistic variability of Moroccan Darija, adversarial training resulted in a notable improvement in both overall model performance and robustness against adversarial attacks, yielding an average increase of 20.89% in resistance to attacks. Furthermore, this approach enabled GPT-4o Mini to achieve an F1-score of 91%, surpassing the current state-of-the-art performance by 6%. These results highlight the importance of incorporating adversarial approaches in low-resource dialectal settings to effectively address linguistic variability and data scarcity.

Visit

doi.org

Tasks

hate speech detectiontext classification

Languages

Arabic, Algerian SpokenArabic, Moroccan Spoken

Licenses

https://creativecommons.org/licenses/by/4.0/

Similaires

Evaluation of Arabic Large Language Models on Moroccan DialectOffensive Language Dataset for Moroccan Arabic dialectMAOffens: Moroccan Arabic Offensive Language DatasetMoroccan Darija Offensive Language Detection DatasetMoroccan Darija Offensive Language Detection DatasetDhati+: Fine-tuned Large Language Models for Arabic Subjectivity Evaluation

Evaluation of Arabic Large Language Models on Moroccan Dialect

Large Language Models (LLMs) have shown outstanding performance in many Natural Language Processing

Offensive Language Dataset for Moroccan Arabic dialect

This dataset card aims to be a base template for new datasets. It has been generated using this raw

MAOffens: Moroccan Arabic Offensive Language Dataset

Moroccan Darija Offensive Language Detection Dataset

The Moroccan Darija offensive language detection dataset is a human-labeled dataset consisting of a

Moroccan Darija Offensive Language Detection Dataset

Ibrahimi, Anass; Mourhir, Asmaa (2023), “Moroccan Darija Offensive Language Detection Dataset”, Mend

Dhati+: Fine-tuned Large Language Models for Arabic Subjectivity Evaluation

Despite its significance, Arabic, a linguistically rich and morphologically complex language, faces