Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

EMMA-500: Enhancing Massively Multilingual Adaptation of Large Language Models

Domaine:

natural language processing

Type de record:

papermodeldataset
Créateur:
Ji,Li,PaaLin
Hôte:avatar
In this work, we introduce EMMA-500, a large-scale multilingual language model continue-trained on texts across 546 languages designed for enhanced multilingual performance, focusing on improving language coverage for low-resource languages. To facilitate continual pre-training, we compile the MaLA corpus, a comprehensive multilingual dataset enriched with curated datasets across diverse domains. Leveraging this corpus, we conduct extensive continual pre-training of the Llama 2 7B model, resulting in EMMA-500, which demonstrates robust performance across a wide collection of benchmarks, including a comprehensive set of multilingual tasks. Our results highlight the effectiveness of continual pre-training in expanding large language models' language capacity, particularly for underrepresented languages, demonstrating significant gains in cross-lingual transfer, task generalization, and language adaptability. We release the MaLA corpus, EMMA-500 model weights, scripts, and model generations.

Visit

arxiv.org

Tasks

language modelingtransfer learning

Languages

Mala

Tags

Computation and Language

Similaires

Massively Multilingual Adaptation of Large Language Models Using Bilingual Translation DataMaLA-500: Massive Language Adaptation of Large Language ModelsGlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language ModelsRomanization-based Large-scale Adaptation of Multilingual Language ModelsSERENGETI: Massively Multilingual Language Models for AfricaOn the Calibration of Massively Multilingual Language Models

Massively Multilingual Adaptation of Large Language Models Using Bilingual Translation Data

This paper investigates a critical design decision in the practice of massively multilingual continu

MaLA-500: Massive Language Adaptation of Large Language Models

Large language models (LLMs) have advanced the state of the art in natural language processing. Howe

GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models

Large language models (LLMs) are advancing at an unprecedented pace globally, with regions increasin

Romanization-based Large-scale Adaptation of Multilingual Language Models

Large multilingual pretrained language models (mPLMs) have become the de facto state of the art for

SERENGETI: Massively Multilingual Language Models for Africa

Multilingual language models (MLMs) acquire valuable, generalizable linguistic information during pretraining and have advanced the state of the art on task-specific finetuning. So far, only ~ 28 out of ~2,000 African languages are covered in existing language mode

On the Calibration of Massively Multilingual Language Models

Massively Multilingual Language Models (MMLMs) have recently gained popularity due to their surprisi