Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Data Augmentation With Back translation for Low Resource languages: A case of English and Luganda

Domaine:

natural language processing

Type de record:

paper
Créateur:
KimHeoRimCho
Hôte:avatar
In this paper,we explore the application of Back translation (BT) as a semi-supervised technique to enhance Neural Machine Translation(NMT) models for the English-Luganda language pair, specifically addressing the challenges faced by low-resource languages. The purpose of our study is to demonstrate how BT can mitigate the scarcity of bilingual data by generating synthetic data from monolingual corpora. Our methodology involves developing custom NMT models using both publicly available and web-crawled data, and applying Iterative and Incremental Back translation techniques. We strategically select datasets for incremental back translation across multiple small datasets, which is a novel element of our approach. The results of our study show significant improvements, with translation performance for the English-Luganda pair exceeding previous benchmarks by more than 10 BLEU score units across all translation directions. Additionally, our evaluation incorporates comprehensive assessment metrics such as SacreBLEU, ChrF2, and TER, providing a nuanced understanding of translation quality. The conclusion drawn from our research confirms the efficacy of BT when strategically curated datasets are utilized, establishing new performance benchmarks and demonstrating the potential of BT in enhancing NMT models for low-resource languages. NLPIR '24: Proceedings of the 2024 8th International Conference on Natural Language Processing and Information Retrieval

Visit

arxiv.org

Tasks

machine translation

Languages

Ganda

Tags

Computation and Language

Similaires

Data Augmentation for Low Resource Neural Machine Translation for Sotho-Tswana LanguagesData Augmentation for Low-Resource Neural Machine TranslationTowards Guided Back-translation for Low-resource languages- A Case Study on Kabyle-FrenchSynthetic Data Diversity vs. Back-Translation for Multilingual NER in Low-Resource LanguagesA Diverse Data Augmentation Strategy for Low-Resource Neural Machine TranslationA Systematic Review of Transfer Learning and Data Augmentation in Neural Machine Translation of Low-Resource Languages

Data Augmentation for Low Resource Neural Machine Translation for Sotho-Tswana Languages

Neural Machine Translation (NMT) models have achieved remarkable performance on translating

Data Augmentation for Low-Resource Neural Machine Translation

The quality of a Neural Machine Translation system depends substantially on the availability of siza

Towards Guided Back-translation for Low-resource languages- A Case Study on Kabyle-French

Synthetic Data Diversity vs. Back-Translation for Multilingual NER in Low-Resource Languages

Named Entity Recognition(NER) for low-resource languages aims to produce robust systems for language

A Diverse Data Augmentation Strategy for Low-Resource Neural Machine Translation

One important issue that affects the performance of neural machine translation is the scale of avail

A Systematic Review of Transfer Learning and Data Augmentation in Neural Machine Translation of Low-Resource Languages

Modern Neural Machine Translation (NMT) systems have achieved state-of-the-art, performance, largely