Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Massively Multilingual Adaptation of Large Language Models Using Bilingual Translation Data

Domaine:

natural language processing

Type de record:

paperdatasetmodel
Créateur:
Ji,Li,PaaLuo
Hôte:avatar
This paper investigates a critical design decision in the practice of massively multilingual continual pre-training -- the inclusion of parallel data. Specifically, we study the impact of bilingual translation data for massively multilingual language adaptation of the Llama3 family of models to 500 languages. To this end, we construct the MaLA bilingual translation corpus, containing data from more than 2,500 language pairs. Subsequently, we develop the EMMA-500 Llama 3 suite of four massively multilingual models -- continually pre-trained from the Llama 3 family of base models extensively on diverse data mixes up to 671B tokens -- and explore the effect of continual pre-training with or without bilingual translation data. Comprehensive evaluation across 7 tasks and 12 benchmarks demonstrates that bilingual data tends to enhance language transfer and performance, particularly for low-resource languages. We open-source the MaLA corpus, EMMA-500 Llama 3 suite artefacts, code, and model generations. EMMA-500 Gen 2; refer to Gen 1 in arXiv:2409.17892

Visit

arxiv.org

Languages

Mala

Tags

Computation and Language

Similaires

EMMA-500: Enhancing Massively Multilingual Adaptation of Large Language ModelsGlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language ModelsRomanization-based Large-scale Adaptation of Multilingual Language ModelsSERENGETI: Massively Multilingual Language Models for AfricaOn the Calibration of Massively Multilingual Language ModelsSocially Responsible Data for Large Multilingual Language Models

EMMA-500: Enhancing Massively Multilingual Adaptation of Large Language Models

In this work, we introduce EMMA-500, a large-scale multilingual language model continue-trained on t

GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models

Large language models (LLMs) are advancing at an unprecedented pace globally, with regions increasin

Romanization-based Large-scale Adaptation of Multilingual Language Models

Large multilingual pretrained language models (mPLMs) have become the de facto state of the art for

SERENGETI: Massively Multilingual Language Models for Africa

Multilingual language models (MLMs) acquire valuable, generalizable linguistic information during pretraining and have advanced the state of the art on task-specific finetuning. So far, only ~ 28 out of ~2,000 African languages are covered in existing language mode

On the Calibration of Massively Multilingual Language Models

Massively Multilingual Language Models (MMLMs) have recently gained popularity due to their surprisi

Socially Responsible Data for Large Multilingual Language Models

Large Language Models (LLMs) have rapidly increased in size and apparent capabilities in the last th