Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

An Exploration of Data Augmentation Techniques for Improving English to Tigrinya Translation

Domaine:

natural language processing

Type de record:

paper
Créateur:
KidKumTsv
Hôte:avatar
It has been shown that the performance of neural machine translation (NMT) drops starkly in low-resource conditions, often requiring large amounts of auxiliary data to achieve competitive results. An effective method of generating auxiliary data is back-translation of target language sentences. In this work, we present a case study of Tigrinya where we investigate several back-translation methods to generate synthetic source sentences. We find that in low-resource conditions, back-translation by pivoting through a higher-resource language related to the target language proves most effective resulting in substantial improvements over baselines. Accepted at AfricaNLP Workshop, EACL 2021

Visit

arxiv.org

Tasks

machine translation

Languages

Tigrigna

Tags

Computation and Language

Similaires

Offline Corpus Augmentation for English-Amharic Machine TranslationTextual Augmentation Techniques Applied to Low Resource Machine Translation: Case of SwahiliData Augmentation With Back translation for Low Resource languages: A case of English and LugandaData Augmentation for Low-Resource Neural Machine TranslationAssessment of Amharic-English Literary Translation Quality Through Data Compression TechniquesImproving low-resource neural machine translation by semantic distance augmentation

Offline Corpus Augmentation for English-Amharic Machine Translation

The purpose of this study was to investigate the effect of corpus augmentation on the quality of English-Amharic Machine Translation (MT). In fact, trigram and four-gram Statistical Machine Translation (SMT) language models, as well as Neural Machine Translation (N

Textual Augmentation Techniques Applied to Low Resource Machine Translation: Case of Swahili

In this work we investigate the impact of applying textual data augmentation tasks to low resource machine translation. There has been recent interest in investigating approaches for training systems for languages with limited resources and one popular approach is

Data Augmentation With Back translation for Low Resource languages: A case of English and Luganda

In this paper,we explore the application of Back translation (BT) as a semi-supervised technique to

Data Augmentation for Low-Resource Neural Machine Translation

The quality of a Neural Machine Translation system depends substantially on the availability of siza

Assessment of Amharic-English Literary Translation Quality Through Data Compression Techniques

Translation quality assessment are crucial challenges in computational linguistics. This study explo

Improving low-resource neural machine translation by semantic distance augmentation

Neural machine translation (NMT) has witnessed substantial advancements, leveraging its learning cap