Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Unlocking the Potential of Model Merging for Low-Resource Languages

Domain:

natural language processing

Record type:

paper
Creator:
TaoZhaHuaMa,
Host:avatar
Adapting large language models (LLMs) to new languages typically involves continual pre-training (CT) followed by supervised fine-tuning (SFT). However, this CT-then-SFT approach struggles with limited data in the context of low-resource languages, failing to balance language modeling and task-solving capabilities. We thus propose model merging as an alternative for low-resource languages, combining models with distinct capabilities into a single model without additional training. We use model merging to develop task-solving LLMs for low-resource languages without SFT data in the target languages. Our experiments based on Llama-2-7B demonstrate that model merging effectively endows LLMs for low-resource languages with task-solving abilities, outperforming CT-then-SFT in scenarios with extremely scarce data. Observing performance saturation in model merging with more training tokens, we further analyze the merging process and introduce a slack variable to the model merging algorithm to mitigate the loss of important parameters, thereby enhancing performance. We hope that model merging can benefit more human languages suffering from data scarcity with its higher data efficiency. To appear in EMNLP2024 Findings

Visit

arxiv.org

Tasks

language modelingtransfer learning

Tags

Computation and LanguageArtificial Intelligence

Similar

Unsupervised Language Model Adaptation for Low-Resource LanguagesUnlocking LLM Safeguards for Low-Resource Languages via Reasoning and Alignment with Minimal Training DataBiomedical Machine Translation for Low-Resource Arabic-Script Languages via Cross-Lingual Transfer and LoRA Adapter MergingInkubaLM: A small language model for low-resource African languagesThe eBible Corpus: Data and Model Benchmarks for Bible Translation for Low-Resource LanguagesModel Transfer for Tagging Low-resource Languages using a Bilingual Dictionary

Unsupervised Language Model Adaptation for Low-Resource Languages

This paper introduces a two-way neural machine translation system from Bengali to English and vice v

Unlocking LLM Safeguards for Low-Resource Languages via Reasoning and Alignment with Minimal Training Data

Recent advances in LLMs have enhanced AI capabilities, but also increased the risk posed by maliciou

Biomedical Machine Translation for Low-Resource Arabic-Script Languages via Cross-Lingual Transfer and LoRA Adapter Merging

We present a systematic study of healthcare-domain cross-lingual transfer to address the scarcity of

InkubaLM: A small language model for low-resource African languages

High-resource language models often fall short in the African context, where there is a critical nee

The eBible Corpus: Data and Model Benchmarks for Bible Translation for Low-Resource Languages

Efficiently and accurately translating a corpus into a low-resource language remains a challenge, re

Model Transfer for Tagging Low-resource Languages using a Bilingual Dictionary

Cross-lingual model transfer is a compelling and popular method for predicting annotations in a low-