Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Can Multilingual Language Models Transfer to an Unseen Dialect? A Case Study on North African Arabizi

Domaine:

natural language processing

Type de record:

paper
Créateur:
MulSagSeddah, Djamé
Éditeur:
arXiv
Hôte:avatar
Building natural language processing systems for non standardized and low resource languages is a difficult challenge. The recent success of large-scale multilingual pretrained language models provides new modeling tools to tackle this. In this work, we study the ability of multilingual language models to process an unseen dialect. We take user generated North-African Arabic as our case study, a resource-poor dialectal variety of Arabic with frequent code-mixing with French and written in Arabizi, a non-standardized transliteration of Arabic to Latin script. Focusing on two tasks, part-of-speech tagging and dependency parsing, we show in zero-shot and unsupervised adaptation scenarios that multilingual language models are able to transfer to such an unseen dialect, specifically in two extreme cases: (i) across scripts, using Modern Standard Arabic as a source language, and (ii) from a distantly related language, unseen during pretraining, namely Maltese. Our results constitute the first successful transfer experiments on this dialect, paving thus the way for the development of an NLP ecosystem for resource-scarce, non-standardized and highly variable vernacular languages.

Visit

doi.orgarxiv.org

Tasks

dependency parsingparsingpart of speech taggingtransfer learning

Tags

Computation and Language (cs.CL)Machine Learning (cs.LG)FOS: Computer and information sciencesFOS: Computer and information sciences

Licenses

Creative Commons Attribution Share Alike 4.0 Internationalhttps://creativecommons.org/licenses/by-sa/4.0/legalcode

Similaires

Unsupervised Learning for Handling Code-Mixed Data: A Case Study on POS Tagging of North-African Arabizi DialectTransfer Learning and Distant Supervision for Multilingual Transformer Models: A Study on African LanguagesSpecializing Multilingual Language Models: An Empirical StudyAn Exploration of Vocabulary Size and Transfer Effects in Multilingual Language Models for African LanguagesLow-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-StudyPretrained self-supervised speech models can recognize unseen consonants

Unsupervised Learning for Handling Code-Mixed Data: A Case Study on POS Tagging of North-African Arabizi Dialect

International audience Language model pretrained representation are now ubiquitous in

Transfer Learning and Distant Supervision for Multilingual Transformer Models: A Study on African Languages

Multilingual transformer models like mBERT and XLM-RoBERTa have obtained great improvements for many NLP tasks on a variety of languages. However, recent works also showed that results from high-resource languages could not be easily transferred to realistic, low-r

Specializing Multilingual Language Models: An Empirical Study

Pretrained multilingual language models have become a common tool in transferring NLP capabilities t

An Exploration of Vocabulary Size and Transfer Effects in Multilingual Language Models for African Languages

Multilingual pretrained language models have been shown to work well on many languages, even those they were not originally pretrained on. Despite their empirical success in downstream tasks, there is still a gap in understanding of "what makes them tick''. In this

Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study

Despite the widespread adoption of Large Language Models (LLMs), their strongest capabilities remain

Pretrained self-supervised speech models can recognize unseen consonants

Modern pretrained self-supervised automatic speech recognition models are trained on large-scale aud