Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

NECAT-CLWE: A Simple But Efficient Parallel Data Generation Approach for Unsupervised and Semi-Supervised Neural Machine Translation

Domain:

natural language processing

Record type:

paper
Many languages lack sufficient data to train qualitative translation systems, particularly those based on the cutting-edge neural machine translation architectures. Recently, it has been demonstrated that using an exact copy of the monolingual target data as the source data improves the quality of translation systems, allowing them to benefit from proper nouns and such similar words that do not require translation. However, using an exact copy of the target data contaminates the source data with terms in the target language that needs translation. As a result, we describe in this paper a similar but more effective parallel data generation approach for improving low-resource neural machine translation using named entity copying and approximate translations using cross-lingual word embedding (NECAT-CLWE). The work will be evaluated on the low resource English-Hausa neural machine translation.

Visit

openreview.net

Tasks

machine translationnamed entity recognition

Languages

Hausa

Tags

africanlp2Parallel Data Generation

Similar

Integrating Unsupervised Data Generation into Self-Supervised Neural Machine Translation for Low-Resource LanguagesTowards Supervised and Unsupervised Neural Machine Translation Baselines for Nigerian PidginSemi-supervised Neural Machine Translation with Consistency Regularization for Low-Resource LanguagesTopXGen: Topic-Diverse Parallel Data Generation for Low-Resource Machine TranslationAmharic-English Parallel Corpus for Neural Machine TranslationA PARALLEL CORPUS FOR ADVANCING ENGLISH–SANTALI NEURAL MACHINE TRANSLATION

Integrating Unsupervised Data Generation into Self-Supervised Neural Machine Translation for Low-Resource Languages

For most language combinations, parallel data is either scarce or simply unavailable. To address thi

Towards Supervised and Unsupervised Neural Machine Translation Baselines for Nigerian Pidgin

Nigerian Pidgin is arguably the most widely spoken language in Nigeria. Variants of this language are also spoken across West and Central Africa, making it a very important language. This work aims to establish supervised and unsupervised neural machine translation

Semi-supervised Neural Machine Translation with Consistency Regularization for Low-Resource Languages

The advent of deep learning has led to a significant gain in machine translation. However, most of t

TopXGen: Topic-Diverse Parallel Data Generation for Low-Resource Machine Translation

LLMs have been shown to perform well in machine translation (MT) with the use of in-context learning

Amharic-English Parallel Corpus for Neural Machine Translation

Amharic is the working language of Ethiopia and, owing to its Semitic characteristics, the language

A PARALLEL CORPUS FOR ADVANCING ENGLISH–SANTALI NEURAL MACHINE TRANSLATION

Machine Translation (MT) poses a significant challenge in developing language corpora for low-resour