Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

ABDUL: a new Approach to Build language models for Dialects Using formal Language corpora only

Domain:

natural language processing

Record type:

papermodel
Creator:
TouSmaLan
Editor:
LabStaACLANR
Publisher:
CCSD
Host:avatar
International audience

Arabic dialects present major challenges for natural language processing (NLP) due to their diglossic nature, phonetic variability, and the scarcity of resources. To address this, we introduce a phoneme-like transcription approach that enables the training of robust language models for North African Dialects (NADs) using only formal language data, without the need for dialect-specific corpora. Our key insight is that Arabic dialects are highly phonetic, with NADs particularly influenced by European languages. This motivated us to develop a novel approach in which we convert Arabic script into a Latin-based representation, allowing our language model, ABDUL, to benefit from existing Latin-script corpora. Our method demonstrates strong performance in multi-label emotion classification and named entity recognition (NER) across various Arabic dialects. ABDUL achieves results comparable to or better than specialized and multilingual models such as DarijaBERT, DziriBERT, and mBERT. Notably, in the NER task, ABDUL outperforms mBERT by 5% in F1-score for Modern Standard Arabic (MSA), Moroccan, and Algerian Arabic, despite using a vocabulary four times smaller than mBERT.

Visit

hal.science

Tasks

emotion identificationinformation extractionlanguage modelingnamed entity recognition

Languages

Arabic, Algerian Spoken

Tags

vocabulary efficiencycomputational linguisticstransfer learninglanguage modelingNorth African Arabic dialectsModern Standard Arabiclow-resource NLP[INFO.INFO-AI]Computer Science [cs]/Artificial Intelligence [cs.AI][INFO]Computer Science [cs][INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL]

Licenses

https://creativecommons.org/licenses/by/4.0/info:eu-repo/semantics/OpenAccess

Similar

Adapting Chat Language Models Using Only Target Unlabeled Language DataA Survey of Large Language Models for Arabic Language and its DialectsUsing language sample analyses across English dialects: A case-based approach for preschoolersSunflower: A New Approach To Expanding Coverage of African Languages in Large Language ModelsSinhala Encoder-only Language Models and EvaluationGlot500: Scaling Multilingual Corpora and Language Models to 500 Languages

Adapting Chat Language Models Using Only Target Unlabeled Language Data

Vocabulary expansion (VE) is the de-facto approach to language adaptation of large language models (

A Survey of Large Language Models for Arabic Language and its Dialects

This survey offers a comprehensive overview of Large Language Models (LLMs) designed for Arabic lang

Using language sample analyses across English dialects: A case-based approach for preschoolers

This study compared language samples from typically developing 4-year-olds who spoke African Amer

Sunflower: A New Approach To Expanding Coverage of African Languages in Large Language Models

There are more than 2000 living languages in Africa, most of which have been bypassed by advances in language technology. Current leading LLMs exhibit strong performance on a number of the most common languages (e.g. Swahili or Yoruba), but prioritise support fo

Sinhala Encoder-only Language Models and Evaluation

Recently, language models (LMs) have produced excellent results in many natural language processing

Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages

The NLP community has mainly focused on scaling Large Language Models (LLMs) vertically, i.e., makin