Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Beyond Many-Shot Translation: Scaling In-Context Demonstrations For Low-Resource Machine Translation

Domaine:

natural language processing

Type de record:

paper
Créateur:
SalCarMorAi,
Hôte:avatar
Building machine translation (MT) systems for low-resource languages is notably difficult due to the scarcity of high-quality data. Although Large Language Models (LLMs) have improved MT system performance, adapting them to lesser-represented languages remains challenging. In-context learning (ICL) may offer novel ways to adapt LLMs for low-resource MT by conditioning models on demonstration at inference time. In this study, we explore scaling low-resource machine translation ICL beyond the few-shot setting to thousands of examples with long-context models. We scale in-context token budget to 1M tokens and compare three types of training corpora used as in-context supervision: monolingual unsupervised data, instruction-style data, and parallel data (English--target and Indonesian--target). Our experiments on Javanese and Sundanese show that gains from additional context saturate quickly and can degrade near the maximum context window, with scaling behavior strongly dependent on corpus type. Notably, some forms of monolingual supervision can be competitive with parallel data, despite the latter offering additional supervision. Overall, our results characterize the effective limits and corpus-type sensitivity of long-context ICL for low-resource MT, highlighting that larger context windows do not necessarily yield proportional quality gains. 8 pages, 18 figures, EACL 2026 Conference - LoResMT workshop

Visit

arxiv.org

Tasks

machine translation

Tags

Computation and Language68T50I.2.7

Similaires

An Empirical Study of Many-Shot In-Context Learning for Machine Translation of Low-Resource LanguagesScaling Zero-shot Machine Translation on Low-Resource Languages with Linguistic GrammarsTranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine TranslationLesan -- Machine Translation for Low Resource LanguagesLesan: Machine Translation for Low Resource LanguagesLow-Resource Machine Translation for Moroccan Arabic

An Empirical Study of Many-Shot In-Context Learning for Machine Translation of Low-Resource Languages

In-context learning (ICL) allows large language models (LLMs) to adapt to new tasks from a few examp

Scaling Zero-shot Machine Translation on Low-Resource Languages with Linguistic Grammars

Most of the world’s languages lack parallel corpora, monolingual web data, or sufficient representat

TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine Translation

The rapid progress in Artificial Intelligence has largely bypassed African languages, creating a dig

Lesan -- Machine Translation for Low Resource Languages

Millions of people around the world can not access content on the Web because most of the content is not readily available in their language. Machine translation (MT) systems have the potential to change this for many languages. Current MT systems provide very accu

Lesan: Machine Translation for Low Resource Languages

Human evaluation dataset to evaluate machine translation systems to and from Amharic, English and Tigrinya.

Low-Resource Machine Translation for Moroccan Arabic