Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Data Augmentation for Low Resource Neural Machine Translation for Sotho-Tswana Languages

Domain:

natural language processing

Record type:

paperdatasetmodel
Creator:
Mojapelo, MaxwellBuys, JanMojapelo, Maxwell
Editor:
Pillay, Anban W.Jembere, EdgarGerber, Aurona
Publisher:
Steering Committee of the Southern African Conference for Artificial Intelligence Research, South Africa
Host:avatar

Neural Machine Translation (NMT) models have achieved remarkable performance on translating between high resource languages. However, translation quality for languages with limited data is much worse. This research focuses on the low resource language of Sepedi and considers two data augmentation techniques to increase the size and diversity of English-Sepedi corpora for training an NMT model. First we consider backtranslation, which makes use of the larger amount of available monolingual Sepedi text. We train a reverse (Sepedi to English) model and generate synthetic English sentences from the monolingual Sepedi sentences. These synthetic translations examples are added to the parallel English-Sepedi sentences. We carry out various experiments to investigate translation quality improvements. The second technique we consider is to generate synthetic data from parallel sentences between English and a closely-related language, Setswana. Setwana word are replacing with Sepedi words through an induced bilingual dictionary, which is created by using a supervised Generative Adversarial Network to align the embeddings of Sepedi and Setswana words. We evaluate our models on the JW300, FLoRes and Autshumato evaluation test sets, finding improvements over the current benchmark BLEU scores across all three datasets.

Visit

doi.org

Tasks

machine translation

Languages

BirwaSetswanaSotho, Northern

Tags

Neural Machine TranslationData AugmentationBacktranslationWord Replacement.

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution Share Alike 4.0 Internationalhttps://creativecommons.org/licenses/by-sa/4.0/legalcode

Similar

Data Augmentation for Low-Resource Neural Machine TranslationA Diverse Data Augmentation Strategy for Low-Resource Neural Machine TranslationMultilingual Neural Machine Translation for Low Resource LanguagesLow-Resource Neural Machine Translation for Southern African LanguagesNeural Machine Translation for Low-Resource Languages: A SurveySelecting data for multilingual multi-domain neural machine translation on low resource languages

Data Augmentation for Low-Resource Neural Machine Translation

The quality of a Neural Machine Translation system depends substantially on the availability of siza

A Diverse Data Augmentation Strategy for Low-Resource Neural Machine Translation

One important issue that affects the performance of neural machine translation is the scale of avail

Multilingual Neural Machine Translation for Low Resource Languages

Neural Machine Translation (NMT) has been shown to be more effective in translation tasks compared t

Low-Resource Neural Machine Translation for Southern African Languages

Low-resource African languages have not fully benefited from the progress in neural machine translat

Neural Machine Translation for Low-Resource Languages: A Survey

Neural Machine Translation (NMT) has seen a tremendous spurt of growth in less than ten years, and h

Selecting data for multilingual multi-domain neural machine translation on low resource languages

[ACCESS RESTRICTED TO THE UNIVERSITY OF MISSOURI AT REQUEST OF AUTHOR.] While machine translation ha