Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Neural Machine Translation Models with Back-Translation for the Extremely Low-Resource Indigenous Language Bribri

Domaine:

natural language processing

Type de record:

datasetmodel
Créateur:
IntCotFel
Éditeur:
Und
Hôte:avatar
This paper presents a neural machine translation model and dataset for the Chibchan language Bribri, with an average performance of BLEU 16.9±1.7. This was trained on an extremely small dataset (5923 Bribri-Spanish pairs), providing evidence for the applicability of NMT in extremely low-resource environments. We discuss the challenges entailed in managing training input from languages without standard orthographies, we provide evidence of successful learning of Bribri grammar, and also examine the translations of structures that are infrequent in major Indo-European languages, such as positional verbs, ergative markers, numerical classifiers and complex demonstrative systems. In addition to this, we perform an experiment of augmenting the dataset through iterative back-translation (Sennrich et al., 2016a; Hoang et al., 2018) by using Spanish sentences to create synthetic Bribri sentences. This improves the score by an average of 1.0 BLEU, but only when the new Spanish sentences belong to the same domain as the other Spanish examples. This contributes to the small but growing body of research on Chibchan NLP.

Visit

doi.orgunderline.io

Tasks

machine translation

Tags

Computer and Information ScienceNatural Language ProcessingNeural Network

Similaires

Continual Mixed-Language Pre-Training for Extremely Low-Resource Neural Machine TranslationNeural Machine Translation for Mooré, a Low-Resource LanguageNeural Machine Translation for Extremely Low-Resource African Languages: A Case Study on BambaraAddressing word-order Divergence in Multilingual Neural Machine Translation for extremely Low Resource LanguagesInteractive Machine Translation with Large Language Models for Low-resource LanguagesLow-resource neural machine translation with morphological modeling

Continual Mixed-Language Pre-Training for Extremely Low-Resource Neural Machine Translation

The data scarcity in low-resource languages has become a bottleneck to building robust neural machin

Neural Machine Translation for Mooré, a Low-Resource Language

Neural Machine Translation for Extremely Low-Resource African Languages: A Case Study on Bambara

Low-resource languages present unique challenges to (neural) machine translation. We discuss the cas

Addressing word-order Divergence in Multilingual Neural Machine Translation for extremely Low Resource Languages

Transfer learning approaches for Neural Machine Translation (NMT) train a NMT model on the assisting

Interactive Machine Translation with Large Language Models for Low-resource Languages

Large language models (LLM) have been applied to machine translation with notable success. However,

Low-resource neural machine translation with morphological modeling

Morphological modeling in neural machine translation (NMT) is a promising approach to achieving open