Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Continual Mixed-Language Pre-Training for Extremely Low-Resource Neural Machine Translation

Domaine:

natural language processing

Type de record:

paper
Créateur:
LiuWinFun
Hôte:avatar
The data scarcity in low-resource languages has become a bottleneck to building robust neural machine translation systems. Fine-tuning a multilingual pre-trained model (e.g., mBART (Liu et al., 2020)) on the translation task is a good approach for low-resource languages; however, its performance will be greatly limited when there are unseen languages in the translation pairs. In this paper, we present a continual pre-training (CPT) framework on mBART to effectively adapt it to unseen languages. We first construct noisy mixed-language text from the monolingual corpus of the target language in the translation pair to cover both the source and target languages, and then, we continue pre-training mBART to reconstruct the original monolingual text. Results show that our method can consistently improve the fine-tuning performance upon the mBART baseline, as well as other strong baselines, across all tested low-resource translation pairs containing unseen languages. Furthermore, our approach also boosts the performance on translation pairs where both languages are seen in the original mBART's pre-training. The code is available at github.com. Accepted in Findings of ACL 2021

Visit

arxiv.org

Tasks

machine translation

Tags

Computation and LanguageArtificial Intelligence

Similaires

Pre-Training on Mixed Data for Low-Resource Neural Machine TranslationNeural Machine Translation Models with Back-Translation for the Extremely Low-Resource Indigenous Language BribriNeural Machine Translation for Extremely Low-Resource African Languages: A Case Study on BambaraAddressing word-order Divergence in Multilingual Neural Machine Translation for extremely Low Resource LanguagesNeural Machine Translation for Mooré, a Low-Resource LanguageLanguage-Family Adapters for Low-Resource Multilingual Neural Machine Translation

Pre-Training on Mixed Data for Low-Resource Neural Machine Translation

The pre-training fine-tuning mode has been shown to be effective for low resource neural machine tra

Neural Machine Translation Models with Back-Translation for the Extremely Low-Resource Indigenous Language Bribri

This paper presents a neural machine translation model and dataset for the Chibchan language Bribri,

Neural Machine Translation for Extremely Low-Resource African Languages: A Case Study on Bambara

Low-resource languages present unique challenges to (neural) machine translation. We discuss the cas

Addressing word-order Divergence in Multilingual Neural Machine Translation for extremely Low Resource Languages

Transfer learning approaches for Neural Machine Translation (NMT) train a NMT model on the assisting

Neural Machine Translation for Mooré, a Low-Resource Language

Language-Family Adapters for Low-Resource Multilingual Neural Machine Translation

Large multilingual models trained with self-supervision achieve state-of-the-art results in a wide r