Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Aspects of Terminological and Named Entity Knowledge within Rule-Based Machine Translation Models for Under-Resourced Neural Machine Translation Scenarios

Domaine:

natural language processing

Type de record:

paper
Créateur:
TorPasMasCha
Hôte:avatar
Rule-based machine translation is a machine translation paradigm where linguistic knowledge is encoded by an expert in the form of rules that translate text from source to target language. While this approach grants extensive control over the output of the system, the cost of formalising the needed linguistic knowledge is much higher than training a corpus-based system, where a machine learning approach is used to automatically learn to translate from examples. In this paper, we describe different approaches to leverage the information contained in rule-based machine translation systems to improve a corpus-based one, namely, a neural machine translation model, with a focus on a low-resource scenario. Three different kinds of information were used: morphological information, named entities and terminology. In addition to evaluating the general performance of the system, we systematically analysed the performance of the proposed approaches when dealing with the targeted phenomena. Our results suggest that the proposed models have limited ability to learn from external information, and most approaches do not significantly alter the results of the automatic evaluation, but our preliminary qualitative evaluation shows that in certain cases the hypothesis generated by our system exhibit favourable behaviour such as keeping the use of passive voice.

Visit

arxiv.org

Tasks

machine translation

Tags

Computation and Language

Similaires

Low Resourced Multilingual Neural Machine Translation for Ometo-EnglishNeural Machine Translation in Low-Resourced Languages: Case of Dholuo-Swahili TranslationBeyond MLE: Investigating SEARNN for Low-Resourced Neural Machine TranslationUnderstanding the effects of word-level linguistic annotations in under-resourced neural machine translationNeural Machine Translation for Amharic-English TranslationCrowdsourced Phrase-Based Tokenization for Low-Resourced Neural Machine Translation: The Case of Fon Language

Low Resourced Multilingual Neural Machine Translation for Ometo-English

In this paper, we present a new approach to overcome the problem of language resources that share si

Neural Machine Translation in Low-Resourced Languages: Case of Dholuo-Swahili Translation

Beyond MLE: Investigating SEARNN for Low-Resourced Neural Machine Translation

Structured prediction tasks, like machine translation, involve learning functions that map structure

Understanding the effects of word-level linguistic annotations in under-resourced neural machine translation

This paper studies the effects of word-level linguistic annotations in under-resourced neural machin

Neural Machine Translation for Amharic-English Translation

This paper describes neural machine translation between orthographically and morphologically divergent languages. Amharic has a rich morphology; it uses the syllabic Ethiopic script. We used a new transliteration technique for Amharic to facilitate vocabulary shari

Crowdsourced Phrase-Based Tokenization for Low-Resourced Neural Machine Translation: The Case of Fon Language

Building effective neural machine translation (NMT) models for very low-resourced and morphologically rich African indigenous languages is an open challenge. Besides the issue of finding available resources for them, a lot of work is put into preprocessing and toke