Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Integrating Syntactic Structures for Enhanced English-Amharic Machine Translation: A Graph2Seq-Based Approach

Domaine:

natural language processing

Type de record:

paper
Créateur:
DepMulShiMes
Éditeur:
In
Hôte:
Background/Objective: The shift from classical rule-based and statistical systems of Machine Translation (MT) to more advanced Neural Machine Translation (NMT) systems employing transformers is quite astonishing. Still, translating language pairs such as English and Amharic, which are structurally different, poses a considerable challenge. This study explores the impact of incorporating known syntactic information on the performance of English-Amharic translation in Graph2Seq model to the precision and fluency of translations. Methods: We aimed to enhance translation performance with the design of an attention-based Graph2Seq model that utilizes a shaped syntactic dependency tree and implements Graph Neural Networks (GNN) to capture the hierarchy of language. The model overcomes the constraints posed by conventional sequence-to-sequence (Seq2Seq) models by adding syntactic information into the encoder. The dataset includes roughly 1.14 million parallel sentences in English and Amharic. Findings: There is a positive, actionable gap, as the proposed model was able to exceed the benchmarks set by the transformer models and pretrained methodologies, attaining BLEU 37.3, and demonstrating marked improvement in translation quality. Furthermore, these findings have defied the argument stating that diversity on the morpho-syntactic level impedes translation, instead illustrating the potency of rich superset syntax modeling. This score beats baseline figures by 4.56 points and 24.24 points compared to M2M100 and the standard transformer models, respectively. Novelty: All these points towards the fact that incorporating deep syntactic structures to enrich Graph-NMT systems on low-resource languages enhances their capabilities, paving avenues for further investigations in the field of Machine Translation. This study can be a starting point towards more research on the enhancement of accuracy and fluency of translations using a multidisciplinary approach based on language. Keywords: Graph neural networks, English-Amharic machine translation, Syntactic of source language, BLEU score, Low-resource language

Visit

doi.org

Tasks

machine translation

Languages

Amharic

Similaires

Enhancing English to Amharic machine translation with prior knowledge integration: Leveraging syntactic structures of the source languagePhoneme-based English-Amharic Statistical Machine TranslationNeural Machine Translation for Amharic-English TranslationDevelopment of a Transformer-Based Model for Amharic to English Neural Machine TranslationEnglish-Amharic Statistical Machine TranslationAmharic-English-Machine-Translation-Corpus

Enhancing English to Amharic machine translation with prior knowledge integration: Leveraging syntactic structures of the source language

Abstract Machine translation has made significant progress in automating the conversion of

Phoneme-based English-Amharic Statistical Machine Translation

International audience

Neural Machine Translation for Amharic-English Translation

This paper describes neural machine translation between orthographically and morphologically divergent languages. Amharic has a rich morphology; it uses the syllabic Ethiopic script. We used a new transliteration technique for Amharic to facilitate vocabulary shari

Development of a Transformer-Based Model for Amharic to English Neural Machine Translation

English-Amharic Statistical Machine Translation

International audience no abstract

Amharic-English-Machine-Translation-Corpus

Amharic English Machine Translation Corpus prepared through website crawelling and custom preprocessing. This is a corpus made in effort to make amaharic english parallel data avilable for anyone who wants to deal with machine translation.