Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Fixing Rogue Memorization in Many-to-One Multilingual Translators of Extremely-Low-Resource Languages by Rephrasing Training Samples

Domaine:

natural language processing

Type de record:

paper
Créateur:
AssCavHenNog
Éditeur:
Und
Hôte:avatar
In this paper we study the fine-tuning of pre-trained large high-resource language models (LLMs) into many-to-one multilingual machine translators for extremely-low-resource languages such as endangered Indigenous languages. We explore those issues using datasets created from pseudo-parallel translations to English of The Bible} written in 39~Brazilian Indigenous languages using mBART50 and WMT19 as pre-trained models and multiple translation metrics. We examine bilingual and multilingual models and show that, according to machine translation metrics, same-linguistic family models tend to perform best. However, we also found that many-to-one multilingual systems have a tendency to learn a “rogue” strategy of storing output strings from the training data in the LLM structure and retrieving them instead of performing actual translations. We show that rephrasing the output of the training samples seems to solve the problem.

Visit

doi.orgunderline.io

Tasks

machine translation

Tags

Computational LinguisticsNatural Language ProcessingArtificial IntelligenceMachine translation

Similaires

Multilingual unsupervised sequence segmentation transfers to extremely low-resource languagesParaKab – Many Languages, One Kabyle: A Multilingual Parallel Corpus for a Low-Resource LanguageFixing MoE Over-Fitting on Low-Resource Languages in Multilingual Machine TranslationEffective vocabulary expansion of multilingual language models for extremely low-resource languagesAddressing word-order Divergence in Multilingual Neural Machine Translation for extremely Low Resource LanguagesMultilingual Intermediate-Task Training for Low-Resource Languages in XTREME

Multilingual unsupervised sequence segmentation transfers to extremely low-resource languages

We show that unsupervised sequence-segmentation performance can be transferred to extremely low-reso

ParaKab – Many Languages, One Kabyle: A Multilingual Parallel Corpus for a Low-Resource Language

Description of the Dataset This dataset consists of three parallel corpora involving the Kabyle lan

Fixing MoE Over-Fitting on Low-Resource Languages in Multilingual Machine Translation

Sparsely gated Mixture of Experts (MoE) models have been shown to be a compute-efficient method to s

Effective vocabulary expansion of multilingual language models for extremely low-resource languages

Multilingual pre-trained language models(mPLMs) offer significant benefits for many low-resource lan

Addressing word-order Divergence in Multilingual Neural Machine Translation for extremely Low Resource Languages

Transfer learning approaches for Neural Machine Translation (NMT) train a NMT model on the assisting

Multilingual Intermediate-Task Training for Low-Resource Languages in XTREME

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni