Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

MoE-CT: A Novel Approach For Large Language Models Training With Resistance To Catastrophic Forgetting

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
Li,Li,XieXio
Hôte:avatar
The advent of large language models (LLMs) has predominantly catered to high-resource languages, leaving a disparity in performance for low-resource languages. Conventional Continual Training (CT) approaches to bridge this gap often undermine a model's original linguistic proficiency when expanding to multilingual contexts. Addressing this issue, we introduce a novel MoE-CT architecture, a paradigm that innovatively separates the base model's learning from the multilingual expansion process. Our design freezes the original LLM parameters, thus safeguarding its performance in high-resource languages, while an appended MoE module, trained on diverse language datasets, augments low-resource language proficiency. Our approach significantly outperforms conventional CT methods, as evidenced by our experiments, which show marked improvements in multilingual benchmarks without sacrificing the model's original language performance. Moreover, our MoE-CT framework demonstrates enhanced resistance to forgetting and superior transfer learning capabilities. By preserving the base model's integrity and focusing on strategic parameter expansion, our methodology advances multilingual language modeling and represents a significant step forward for low-resource language inclusion in LLMs, indicating a fruitful direction for future research in language technologies. 13 pages, 2 figures

Visit

arxiv.org

Tasks

language modelingtransfer learning

Tags

Computation and LanguageArtificial Intelligence

Similaires

Scalable and Efficient MoE Training for Multitask Multilingual ModelsPula: Training Large Language Models for SetswanaATLAS: Efficient Learning Without Catastrophic ForgettingUnveiling Language Routing Isolation in Multilingual MoE Models for Interpretable Subnetwork AdaptationTypological Feature Prediction with Large Language Models: An In-Context Learning ApproachSunflower: A New Approach To Expanding Coverage of African Languages in Large Language Models

Scalable and Efficient MoE Training for Multitask Multilingual Models

The Mixture of Experts (MoE) models are an emerging class of sparsely activated deep learning models

Pula: Training Large Language Models for Setswana

In this work we present Pula, a suite of bilingual language models proficient in both Setswana and E

ATLAS: Efficient Learning Without Catastrophic Forgetting

ATLAS: Efficient Learning Without Catastrophic Forgetting

Poster presented at the Deep Learning Indaba 2022 by Heinrich van Deventer

Unveiling Language Routing Isolation in Multilingual MoE Models for Interpretable Subnetwork Adaptation

Mixture-of-Experts (MoE) models exhibit striking performance disparities across languages, yet the i

Typological Feature Prediction with Large Language Models: An In-Context Learning Approach

Typological features are widely used in multilingual NLP, and the prediction of such features holds

Sunflower: A New Approach To Expanding Coverage of African Languages in Large Language Models

There are more than 2000 living languages in Africa, most of which have been bypassed by advances in language technology. Current leading LLMs exhibit strong performance on a number of the most common languages (e.g. Swahili or Yoruba), but prioritise support fo