Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Embedding Initialization for Unseen Low-resource Languages in Multilingual NMT: A Case Study on Limbum-English Translation

Domain:

natural language processing

Record type:

datasetpaper
Creator:
NtoTuyKafOgo
Publisher:
arXiv
Host:avatar
Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu languages of Cameroon. When fine-tuning these models for an unseen language, practitioners must choose a proxy language token, yet no principled method exists for this selection. We implemented an embedding initialization strategy where a language token is the average of embeddings from multiple typologically related languages already in the mod el. We evaluate this approach on Limbum-to-English translation using a parallel corpus of 8,837 sentence pairs from New Testament text and a bilingual dictionary. We compare models: NLLB-200 zero-shot (chrF2++ = 12.5), a Transformer trained from scratch (chrF2++ = 14.5), NLLB-200 fine-tuned with a Swahili proxy token (chrF2++ = 47.3), and NLLB-200 with our averaged embedding initialization (chrF2++ = 46.7). We find that the multi-language initialization achieves performance comparable to the best single-language proxy. Both NLLB-200 variants improve over the from-scratch baseline by over 32 chrF2++ points. These results show that multilingual transfer is the dominant factor in extremely low-resource Bantu translation while eliminating the need for heuristic proxy selection. However, all systems fail to preserve tonal diacritics, highlighting an open challenge. We make our dataset and code available to support further research.

Visit

doi.org

Tasks

machine translation

Languages

LimbumSwahili

Tags

Computation and Language (cs.CL)Machine Learning (cs.LG)FOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

LGSE: Lexically Grounded Subword Embedding Initialization for Low-Resource Language AdaptationError Analysis of Multilingual Language Models in Machine Translation for Low-resource Languages: A Case Study of Amharic to English Bi-directional Machine TranslationMultilingual Intermediate-Task Training for Robust NMT in Low-Resource Languages on the XTREME BenchmarkMassively Multilingual Text Translation For Low-Resource LanguagesMultilingual Neural Machine Translation for Low Resource LanguagesNeural Machine Translation for Extremely Low-Resource African Languages: A Case Study on Bambara

LGSE: Lexically Grounded Subword Embedding Initialization for Low-Resource Language Adaptation

Adapting pretrained language models to low-resource, morphologically rich languages remains a signif

Error Analysis of Multilingual Language Models in Machine Translation for Low-resource Languages: A Case Study of Amharic to English Bi-directional Machine Translation

Multilingual large language models (mLLMs) have significantly advanced machine translation, yet chal

Multilingual Intermediate-Task Training for Robust NMT in Low-Resource Languages on the XTREME Benchmark

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Massively Multilingual Text Translation For Low-Resource Languages

Translation into severely low-resource languages has both the cultural goal of saving and reviving t

Multilingual Neural Machine Translation for Low Resource Languages

Neural Machine Translation (NMT) has been shown to be more effective in translation tasks compared t

Neural Machine Translation for Extremely Low-Resource African Languages: A Case Study on Bambara

Low-resource languages present unique challenges to (neural) machine translation. We discuss the cas