Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Scaling Zero-shot Machine Translation on Low-Resource Languages with Linguistic Grammars

Domaine:

natural language processing

Type de record:

paper
Créateur:
EvgValAlbTat
Éditeur:
LCC
Hôte:
Most of the world’s languages lack parallel corpora, monolingual web data, or sufficient representation in multilingual LLM pretraining. However, many are documented in descriptive grammars containing exhaustive syntactic information, interlinear glossed examples and translations. Recent work has shown that large language models can leverage grammars in-context for zero-shot translation and typological classification, but it remains unclear whether grammars alone can serve as the primary supervision for training machine translation systems. We introduce a scalable grammar-centered framework for machine translation, converting descriptive grammars into structured LLM-readable context by extracting example sentences, glosses, and leveraging available external typological metadata. Across five typologically diverse low-resource languages, we present structure-aware regimes that incorporate gloss-level information and grammar conditioning to enable in-context learning for machine translation. We further analyze how translation quality scales with the number of grammar-derived examples – for Georgian (49.01 ChF), Chamorro (53.32 ChF), Basque (63.90 ChF), Igbo (52.83), and Korean (37.09 ChF) languages. Our results show that (1) grammars alone can support non-trivial translation performance when nothing else is available, (2) incorporating typological metadata does not show consistent generalization over flat sentence pairs, and (3) sentence-level examples matter most: extracted parallel examples consistently improve translation quality relative to zero-shot and typology-only prompting. This work reframes descriptive grammars as computational resources rather than static references, offering a practical pathway toward scaling machine translation to languages that have no corpus data. The code for the project will be publicly available at: github.com

Visit

doi.org

Tasks

machine translation

Similaires

Beyond Many-Shot Translation: Scaling In-Context Demonstrations For Low-Resource Machine TranslationCharSpan: Utilizing Lexical Similarity to Enable Zero-Shot Machine Translation for Extremely Low-resource LanguagesSelectNoise: Unsupervised Noise Injection to Enable Zero-Shot Machine Translation for Extremely Low-resource LanguagesImpact of Intermediate Task Scaling on Zero-Shot Cross-Lingual Transfer in Low-Resource LanguagesScaling Model Size and Zero-Shot Transfer to Low-Resource Languages in XTREMEScaling Bilingual Lexicons for Zero-Shot Cross-Lingual Retrieval in Low-Resource Languages

Beyond Many-Shot Translation: Scaling In-Context Demonstrations For Low-Resource Machine Translation

Building machine translation (MT) systems for low-resource languages is notably difficult due to the

CharSpan: Utilizing Lexical Similarity to Enable Zero-Shot Machine Translation for Extremely Low-resource Languages

We address the task of machine translation (MT) from extremely low-resource language (ELRL) to Engli

SelectNoise: Unsupervised Noise Injection to Enable Zero-Shot Machine Translation for Extremely Low-resource Languages

In this work, we focus on the task of machine translation (MT) from extremely low-resource language

Impact of Intermediate Task Scaling on Zero-Shot Cross-Lingual Transfer in Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Scaling Model Size and Zero-Shot Transfer to Low-Resource Languages in XTREME

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Scaling Bilingual Lexicons for Zero-Shot Cross-Lingual Retrieval in Low-Resource Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to