Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

A Generative-Adversarial Approach to Low-Resource Language Translation via Data Augmentation

Domaine:

natural language processing

Type de record:

paper
Créateur:
Lin
Éditeur:
rSc
Hôte:
Language and culture preservation is a serious challenge both socially and technologically. This paper proposes a novel, data augmentation approach to using machine learning to translate low-resource languages. Since low-resource languages, such as Aymara and Quechua, do not have many available translations that machine learning software can use as reference, machine translation models frequently err when translating to and from low-resource languages. Because models learn the syntactic and lexical patterns underlying translations through processing the training data, an insufficient amount of data hinders them from producing accurate translations. In this paper, I propose the novel application of a generative-adversarial network (GAN) to automatically augment low-resource language data. A GAN consists of two competing models, with one learning to generate sentences from noise and the other trying to tell if a given sentence is real or generated. My experiments show that even when training on a very small amount of language data (< 20,000 sentences) in a simulated low-resource setting, such a model is able to generate original, coherent sentences, such as "ask me that healthy lunch im cooking up,” and “my grandfather work harder than your grandfather before.” The first of its kind, this novel application of a GAN is effective in augmenting low-resource language data to improve the accuracy of machine translation and provides a reference for future experimentation with GANs in machine translation.

Visit

doi.org

Tasks

machine translation

Licenses

https://creativecommons.org/licenses/by-nc-sa/4.0

Similaires

Generative Adversarial Networks for Synthetic Data Augmentation in Low-Resource Language Modeling with Cross-Lingual Knowledge TransferCommonsense Knowledge Augmentation for Low-Resource Languages via Adversarial LearningLow-data? No problem: low-resource, language-agnostic conversational text-to-speech via F0-conditioned data augmentationData Augmentation for Low-Resource Neural Machine TranslationData Augmentation via Dependency Tree Morphing for Low-Resource LanguagesProjection-based Data Augmentation and Adversarial Training for Robust Low-resource NER

Generative Adversarial Networks for Synthetic Data Augmentation in Low-Resource Language Modeling with Cross-Lingual Knowledge Transfer

Low-resource language modeling is a challenge addressed in this research using a Generative Adversar

Commonsense Knowledge Augmentation for Low-Resource Languages via Adversarial Learning

Commonsense reasoning is one of the ultimate goals of artificial intelligence research because it si

Low-data? No problem: low-resource, language-agnostic conversational text-to-speech via F0-conditioned data augmentation

The availability of data in expressive styles across languages is limited, and recording sessions ar

Data Augmentation for Low-Resource Neural Machine Translation

The quality of a Neural Machine Translation system depends substantially on the availability of siza

Data Augmentation via Dependency Tree Morphing for Low-Resource Languages

Neural NLP systems achieve high scores in the presence of sizable training dataset. Lack of such dat

Projection-based Data Augmentation and Adversarial Training for Robust Low-resource NER

Named Entity Recognition(NER) for low-resource languages aims to produce robust systems for language