Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Context Volume Drives Performance: Tackling Domain Shift in Extremely Low-Resource Translation via RAG

Domaine:

natural language processing

Type de record:

paper
Créateur:
SetMerLau
Hôte:avatar
Neural Machine Translation (NMT) models for low-resource languages suffer significant performance degradation under domain shift. We quantify this challenge using Dhao, an indigenous language of Eastern Indonesia with no digital footprint beyond the New Testament (NT). When applied to the unseen Old Testament (OT), a standard NMT model fine-tuned on the NT drops from an in-domain score of 36.17 chrF++ to 27.11 chrF++. To recover this loss, we introduce a hybrid framework where a fine-tuned NMT model generates an initial draft, which is then refined by a Large Language Model (LLM) using Retrieval-Augmented Generation (RAG). The final system achieves 35.21 chrF++ (+8.10 recovery), effectively matching the original in-domain quality. Our analysis reveals that this performance is driven primarily by the number of retrieved examples rather than the choice of retrieval algorithm. Qualitative analysis confirms the LLM acts as a robust "safety net," repairing severe failures in zero-shot domains.

Visit

arxiv.org

Tasks

machine translation

Tags

Computation and LanguageArtificial Intelligence

Similaires

In-Context Example Selection via Similarity Search Improves Low-Resource Machine TranslationRobustness of Zero-Shot Cross-Lingual Retrieval Models Against Domain Shift in Low-Resource Languages via ArtificialPredicting Machine Translation Performance on Low-Resource Languages: The Role of Domain SimilarityTranslation or Recitation? Calibrating Evaluation Scores for Machine Translation of Extremely Low-Resource LanguagesIt's All About In-Context Learning! Teaching Extremely Low-Resource Languages to LLMsContinual Mixed-Language Pre-Training for Extremely Low-Resource Neural Machine Translation

In-Context Example Selection via Similarity Search Improves Low-Resource Machine Translation

The ability of generative large language models (LLMs) to perform in-context learning has given rise

Robustness of Zero-Shot Cross-Lingual Retrieval Models Against Domain Shift in Low-Resource Languages via Artificial

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Predicting Machine Translation Performance on Low-Resource Languages: The Role of Domain Similarity

Fine-tuning and testing a multilingual large language model is expensive and challenging for low-res

Translation or Recitation? Calibrating Evaluation Scores for Machine Translation of Extremely Low-Resource Languages

The landscape of extremely low-resource machine translation (MT) is characterized by perplexing vari

It's All About In-Context Learning! Teaching Extremely Low-Resource Languages to LLMs

Extremely low-resource languages, especially those written in rare scripts, as shown in Figure 1, re

Continual Mixed-Language Pre-Training for Extremely Low-Resource Neural Machine Translation

The data scarcity in low-resource languages has become a bottleneck to building robust neural machin