Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Darijabert: a Step Forward in Nlp for the Written Moroccan Dialect

Domain:

natural language processing

Record type:

modeldataset
Creator:
KamAbdou Mohamed NairaAnass AllakImade Benelallam
Publisher:
Res
Host:
Abstract The performance of existing transformer-based language models in providing state-of-the-art results on many downstream tasks is well established. However, these models tend to be limited to high-resource languages or are multilingual in nature. The availability of models dedicated to Arabic dialects is limited, and even those that exist primarily support dialects written in Arabic script. This study presents the first BERT models for Moroccan Arabic dialect, also known as Darija, called DarijaBERT, DarijaBERT-arabizi, and DarijaBERT-mix. These models are trained on the largest Arabic monodialectal corpus, supporting both Arabic and Latin character representations of the Moroccan dialect. The models' performance is evaluated and compared to existing multidialectal and multilingual models on four distinct downstream tasks, demonstrating state-of-the-art results. The data collection methodology and pre-training process are described, and the Moroccan Topic Classification Dataset (MTCD) is introduced as the first dataset for topic classification in the Moroccan Arabic dialect. The pre-trained models and MTCD dataset are available to the scientific community.

Visit

doi.org

Tasks

language modelingtext classificationtopic classification

Languages

Arabic, Algerian SpokenArabic, Moroccan Spoken

Licenses

https://creativecommons.org/licenses/by/4.0/

Similar

MANorm: A Normalization Dictionary for Moroccan Arabic Dialect Written in Latin ScriptMoroccan NLPOne step forward, one step backwards: African regimes' changing relations with artisanal minersMSA-Moroccan Dialect: A Multimodal Sentiment Analysis Dataset for Moroccan Arabic (Darija)Ibom NLP: A Step Toward Inclusive Natural Language Processing for Nigeria's Minority LanguagesDarNERcorp: a Named Entity Recognition Corpus in the Moroccan Dialect

MANorm: A Normalization Dictionary for Moroccan Arabic Dialect Written in Latin Script

Social media user-generated text is actually the main resource for many NLP tasks. This text however

Moroccan NLP

Development and evaluation of NLP models for Moroccan Darija and Arabic, covering code-switching det

One step forward, one step backwards: African regimes' changing relations with artisanal miners

This thesis explores and explains differences in how resource-rich African countries respond to arti

MSA-Moroccan Dialect: A Multimodal Sentiment Analysis Dataset for Moroccan Arabic (Darija)

This dataset provides the first publicly available multimodal resource for sentiment analysis in Mor

Ibom NLP: A Step Toward Inclusive Natural Language Processing for Nigeria's Minority Languages

Nigeria is the most populous country in Africa with a population of more than 200 million people. More than 500 languages are spoken in Nigeria and it is one of the most linguistically diverse countries in the world. Despite this, natural language processing (NL

DarNERcorp: a Named Entity Recognition Corpus in the Moroccan Dialect

DarNERcorp is a manually annotated corpus for Named Entity Recognition (NER) in the Moroccan Dialect