International audience
In a landscape dominated by large language models, Maghrebi Arabic dialects, though widely used in everyday communication and informal writing, remain largely underserved by Natural Language Processing (NLP) technologies. Their limited linguistic resources, high variability, and lack of standardized orthography make them particularly challenging to model effectively. To address these issues, this work introduces MagBERT, a lightweight variant of BERT designed specifically for the three major Maghrebi dialects: Algerian, Moroccan, and Tunisian Arabic, in both Arabic and Latin scripts. The model was pretrained then fine-tuned on multiple downstream tasks, demonstrating competitive performance compared to several strong benchmark models. Despite its compact size, MagBERT shows strong potential as an efficient and versatile model for processing under-resourced North African dialects.