Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Pre-Training on Mixed Data for Low-Resource Neural Machine Translation

Domain:

natural language processing

Record type:

paper
Creator:
WenXiaYatRui
Publisher:
MDP
Host:
The pre-training fine-tuning mode has been shown to be effective for low resource neural machine translation. In this mode, pre-training models trained on monolingual data are used to initiate translation models to transfer knowledge from monolingual data into translation models. In recent years, pre-training models usually take sentences with randomly masked words as input, and are trained by predicting these masked words based on unmasked words. In this paper, we propose a new pre-training method that still predicts masked words, but randomly replaces some of the unmasked words in the input with their translation words in another language. The translation words are from bilingual data, so that the data for pre-training contains both monolingual data and bilingual data. We conduct experiments on Uyghur-Chinese corpus to evaluate our method. The experimental results show that our method can make the pre-training model have a better generalization ability and help the translation model to achieve better performance. Through a word translation task, we also demonstrate that our method enables the embedding of the translation model to acquire more alignment knowledge.

Visit

doi.org

Tasks

machine translation

Licenses

https://creativecommons.org/licenses/by/4.0/

Similar

Continual Mixed-Language Pre-Training for Extremely Low-Resource Neural Machine TranslationData Augmentation for Low-Resource Neural Machine TranslationSelecting data for multilingual multi-domain neural machine translation on low resource languagesA Diverse Data Augmentation Strategy for Low-Resource Neural Machine TranslationPivot pre-finetuning for low-resource machine translationData Augmentation for Low Resource Neural Machine Translation for Sotho-Tswana Languages

Continual Mixed-Language Pre-Training for Extremely Low-Resource Neural Machine Translation

The data scarcity in low-resource languages has become a bottleneck to building robust neural machin

Data Augmentation for Low-Resource Neural Machine Translation

The quality of a Neural Machine Translation system depends substantially on the availability of siza

Selecting data for multilingual multi-domain neural machine translation on low resource languages

[ACCESS RESTRICTED TO THE UNIVERSITY OF MISSOURI AT REQUEST OF AUTHOR.] While machine translation ha

A Diverse Data Augmentation Strategy for Low-Resource Neural Machine Translation

One important issue that affects the performance of neural machine translation is the scale of avail

Pivot pre-finetuning for low-resource machine translation

Pivot pre-finetuning for low-resource machine translation

Poster presented at the Deep Learning Indaba 2023 by Stephen Kiilu

Data Augmentation for Low Resource Neural Machine Translation for Sotho-Tswana Languages

Neural Machine Translation (NMT) models have achieved remarkable performance on translating