Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Adapting Pre-trained Language Models to African Languages via Multilingual Adaptive Fine-Tuning

Domaine:

natural language processing

Type de record:

paper

Multilingual pre-trained language models (PLMs) have demonstrated impressive performance on several downstream tasks for both high-resourced and low-resourced languages. However, there is still a large performance drop for languages unseen during pre-training, especially African languages. One of the most effective approaches to adapt to a new language is language adaptive fine-tuning (LAFT) — fine-tuning a multilingual PLM on monolingual texts of a language using the pre-training objective. However, adapting to target language individually takes large disk space and limits the cross-lingual transfer abilities of the resulting models because they have been specialized for a single language. In this paper, we perform multilingual adaptive fine-tuning on 17 most-resourced African languages and three other high-resource languages widely spoken on the African continent to encourage cross-lingual transfer learning. To further specialize the multilingual PLM, we removed vocabulary tokens from the embedding layer that corresponds to non-African writing scripts before MAFT, thus reducing the model size by around 50%. Our evaluation on two multilingual PLMs (AfriBERTa and XLM-R) and three NLP tasks (NER, news topic classification, and sentiment classification) shows that our approach is competitive to applying LAFT on individual languages while requiring significantly less disk space. Additionally, we show that our adapted PLM also improves the zero-shot cross-lingual transfer abilities of parameter efficient fine-tuning methods.

Visit

aclanthology.orgarxiv

Connected records

model

Tasks

transfer learning

Languages

AmharicDholuoGandaHausaIgboKinyarwandaPidgin, NigerianSwahiliWolofYoruba

Tags

language adapteradaptive finetuning

Similaires

Increasing linguistic diversity in NLP : Fine-tuning Multilingual Pre-trained African Language ModelsMultilingual language model Adaptive Fine-Tuning: A Study on African LanguagesDisfluent-to-Fluent Tunisian Dialect Speech Translation with Fine-Tuning Pre-trained Language ModelsAfroXLMR-Social: Adapting Pre-trained Language Models for African Languages Social Media TextMULTILINGUAL ADAPTIVE FINE-TUNING (MAFT)From N-grams to Pre-trained Multilingual Models For Language Identification

Increasing linguistic diversity in NLP : Fine-tuning Multilingual Pre-trained African Language Models

Increasing linguistic diversity in NLP : Fine-tuning Multilingual Pre-trained African Language Models

Poster presented at the Deep Learning Indaba 2023 by Fiskani Banda

Multilingual language model Adaptive Fine-Tuning: A Study on African Languages

Multilingual pre-trained language models (PLMs) have demonstrated impressive performance on several downstream tasks on both high-resourced and low-resourced languages. However, there is still a large performance drop for languages unseen during pre-training, espec

Disfluent-to-Fluent Tunisian Dialect Speech Translation with Fine-Tuning Pre-trained Language Models

AfroXLMR-Social: Adapting Pre-trained Language Models for African Languages Social Media Text

Language models built from various sources are the foundation of today's NLP progress. However, for

MULTILINGUAL ADAPTIVE FINE-TUNING (MAFT)

We introduce MAFT as an approach to adapt a multi-lingual PLM to a new set of languages. Adapting PLMs has been shown to be effective when adapting to a new domain (Gururangan et al., 2020) or language (Pfeiffer et al., 2020; Alabi et al., 2020; Adelani et al., 202

From N-grams to Pre-trained Multilingual Models For Language Identification

In this paper, we investigate the use of N-gram models and Large Pre-trained Multilingual models for