Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Lifting the Curse of Multilinguality by Pre-training Modular Transformers

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
PfeGoyLinLi,
Hôte:avatar
Multilingual pre-trained models are known to suffer from the curse of multilinguality, which causes per-language performance to drop as they cover more languages. We address this issue by introducing language-specific modules, which allows us to grow the total capacity of the model, while keeping the total number of trainable parameters per language constant. In contrast with prior work that learns language-specific components post-hoc, we pre-train the modules of our Cross-lingual Modular (X-Mod) models from the start. Our experiments on natural language inference, named entity recognition and question answering show that our approach not only mitigates the negative interference between languages, but also enables positive transfer, resulting in improved monolingual and cross-lingual performance. Furthermore, our approach enables adding languages post-hoc with no measurable drop in performance, no longer limiting the model usage to the set of pre-trained languages. NAACL 2022

Visit

arxiv.org

Tags

Computation and Language

Similaires

tylerachang/curse-of-multilingualitySpeculative Decoding and the Curse of MultilingualityOpen-Domain Response Generation in Low-Resource Settings using Self-Supervised Pre-Training of Warm-Started TransformersBreaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech Encoders

tylerachang/curse-of-multilinguality

When is multilinguality a curse? Language modeling for 250 high- and low-resource languages (EMNLP 2

Speculative Decoding and the Curse of Multilinguality

Speculative decoding is a popular technique for large language model (LLM) inference, enabling faste

Open-Domain Response Generation in Low-Resource Settings using Self-Supervised Pre-Training of Warm-Started Transformers

Learning response generation models constitute the main component of building open-domain dialogue s

Breaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech Encoders

Multimodal large language models (MLLMs) have achieved significant success in speech-to-text transla