Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Understanding the effects of word-level linguistic annotations in under-resourced neural machine translation

Domaine:

natural language processing

Type de record:

paper
Créateur:
SánPérSán
Hôte:avatar
This paper studies the effects of word-level linguistic annotations in under-resourced neural machine translation, for which there is incomplete evidence in the literature. The study covers eight language pairs, different training corpus sizes, two architectures, and three types of annotation: dummy tags (with no linguistic information at all), part-of-speech tags, and morpho-syntactic description tags, which consist of part of speech and morphological features. These linguistic annotations are interleaved in the input or output streams as a single tag placed before each word. In order to measure the performance under each scenario, we use automatic evaluation metrics and perform automatic error classification. Our experiments show that, in general, source-language annotations are helpful and morpho-syntactic descriptions outperform part of speech for some language pairs. On the contrary, when words are annotated in the target language, part-of-speech tags systematically outperform morpho-syntactic description tags in terms of automatic evaluation metrics, even though the use of morpho-syntactic description tags improves the grammaticality of the output. We provide a detailed analysis of the reasons behind this result. COLING 2020

Visit

arxiv.org

Tasks

machine translation

Tags

Computation and Language

Similaires

Neural Machine Translation in Low-Resourced Languages: Case of Dholuo-Swahili TranslationAspects of Terminological and Named Entity Knowledge within Rule-Based Machine Translation Models for Under-Resourced Neural Machine Translation ScenariosLow Resourced Multilingual Neural Machine Translation for Ometo-EnglishBeyond MLE: Investigating SEARNN for Low-Resourced Neural Machine TranslationCrowdsourced Phrase-Based Tokenization for Low-Resourced Neural Machine Translation: The Case of Fon LanguageAnalysis and evaluation of comparable corpora for under resourced areas of machine translation

Neural Machine Translation in Low-Resourced Languages: Case of Dholuo-Swahili Translation

Aspects of Terminological and Named Entity Knowledge within Rule-Based Machine Translation Models for Under-Resourced Neural Machine Translation Scenarios

Rule-based machine translation is a machine translation paradigm where linguistic knowledge is encod

Low Resourced Multilingual Neural Machine Translation for Ometo-English

In this paper, we present a new approach to overcome the problem of language resources that share si

Beyond MLE: Investigating SEARNN for Low-Resourced Neural Machine Translation

Structured prediction tasks, like machine translation, involve learning functions that map structure

Crowdsourced Phrase-Based Tokenization for Low-Resourced Neural Machine Translation: The Case of Fon Language

Building effective neural machine translation (NMT) models for very low-resourced and morphologically rich African indigenous languages is an open challenge. Besides the issue of finding available resources for them, a lot of work is put into preprocessing and toke

Analysis and evaluation of comparable corpora for under resourced areas of machine translation