Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

In-Context Example Selection via Similarity Search Improves Low-Resource Machine Translation

Domaine:

natural language processing

Type de record:

paper
Créateur:
ZebSagot, BenoîtBaw
Hôte:avatar
The ability of generative large language models (LLMs) to perform in-context learning has given rise to a large body of research into how best to prompt models for various natural language processing tasks. In this paper, we focus on machine translation (MT), a task that has been shown to benefit from in-context translation examples. However no systematic studies have been published on how best to select examples, and mixed results have been reported on the usefulness of similarity-based selection over random selection. We provide a study covering multiple LLMs and multiple in-context example retrieval strategies, comparing multilingual sentence embeddings. We cover several language directions, representing different levels of language resourcedness (English into French, German, Swahili and Wolof). Contrarily to previously published results, we find that sentence embedding similarity can improve MT, especially for low-resource language directions, and discuss the balance between selection pool diversity and quality. We also highlight potential problems with the evaluation of LLM-based MT and suggest a more appropriate evaluation protocol, adapting the COMET metric to the evaluation of LLMs. Code and outputs are freely available at ArmelRandy/ICL-MT.

Visit

arxiv.org

Tasks

machine translation

Languages

SwahiliWolof

Tags

Computation and Language

Similaires

Filtered Pseudo-parallel Corpus Improves Low-resource Neural Machine TranslationReflective Translation: Improving Low-Resource Machine Translation via Structured Self-ReflectionImproving Low-Resource Machine Translation via Round-Trip Reinforcement LearningPredicting Machine Translation Performance on Low-Resource Languages: The Role of Domain SimilarityBeyond Many-Shot Translation: Scaling In-Context Demonstrations For Low-Resource Machine TranslationUnsupervised Neural Machine Translation for Low-Resource Domains via Meta-Learning

Filtered Pseudo-parallel Corpus Improves Low-resource Neural Machine Translation

Large-scale parallel corpora are essential for training high-quality machine translation systems; ho

Reflective Translation: Improving Low-Resource Machine Translation via Structured Self-Reflection

Low-resource languages such as isiZulu and isiXhosa face persistent challenges in machine translatio

Improving Low-Resource Machine Translation via Round-Trip Reinforcement Learning

Low-resource machine translation (MT) has gained increasing attention as parallel data from low-reso

Predicting Machine Translation Performance on Low-Resource Languages: The Role of Domain Similarity

Fine-tuning and testing a multilingual large language model is expensive and challenging for low-res

Beyond Many-Shot Translation: Scaling In-Context Demonstrations For Low-Resource Machine Translation

Building machine translation (MT) systems for low-resource languages is notably difficult due to the

Unsupervised Neural Machine Translation for Low-Resource Domains via Meta-Learning

Unsupervised machine translation, which utilizes unpaired monolingual corpora as training data, has