Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Standardizing Arabic Dialects for NLP: A BERT-Based Transcoding Approach with a Focus on Moroccan Darija

Domaine:

natural language processing

Type de record:

paper
Créateur:
LTIH. S. LTI
Éditeur:
Lvi
Hôte:
Processing Arabic dialects in Natural Language Processing (NLP) presents significant challenges due to linguistic diversity and the lack of standardized resources. While Modern Standard Arabic (MSA) benefits from advanced NLP tools and extensive annotated datasets, dialects such as Moroccan Darija remain underrepresented. This study introduces a BERT-based transcoding framework that bridges the gap between dialectal Arabic and MSA, enabling the use of pre-trained models optimized for MSA, such as AraBERT. By integrating contextual multilingual embeddings, the proposed approach preserves semantic accuracy while addressing the challenges of dialectal variation. Experimental evaluations on the MAC dataset demonstrate the framework's effectiveness, with the proposed approach significantly outperforming existing models, including DarijaBERT and mBERT, across all key metrics. The findings highlight the scalability of the framework, making it applicable to other Arabic dialects and broader NLP tasks. This research advances Arabic language technology by providing a robust and scalable solution for dialectal NLP, particularly for sentiment analysis and similar downstream applications.

Visit

doi.org

Tasks

text normalization

Languages

Arabic, Algerian SpokenArabic, Moroccan Spoken