The linguistic diversity of Morocco, with Darija and Arabic as its two primary languages, poses significant challenges for Natural Language Processing (NLP) applications. This project develops and evaluates NLP models for Darija and Arabic, covering code-switching detection, Darija language modeling, Arabic-Darija translation, sentiment analysis, and named entity recognition, built on the Moroccan Dialect Corpus (MDC).
The project integrates DarijaBERT (SI2M-Lab/INSEA), the first open-source BERT model for Moroccan Darija, alongside a baseline classifier and curated linguistic resources spanning technology, economy, linguistics, policy, law, education, and health domains. Results demonstrate the feasibility of tailored NLP models for Darija and Arabic, with implications for language teaching, digital literacy, and language-based technology development in Morocco. The package is released open-source under the MIT License, with source and data mirrored across GitHub, GitLab, Bitbucket, Codeberg, and PyPI, and archived with a citable DOI on Zenodo and a preregistered protocol on OSF.