Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Very Low Resource Sentence Alignment: Luhya and Swahili

Domain:

natural language processing

Record type:

paperdataset
Creator:
ChiBas
Host:avatar
Language-agnostic sentence embeddings generated by pre-trained models such as LASER and LaBSE are attractive options for mining large datasets to produce parallel corpora for low-resource machine translation. We test LASER and LaBSE in extracting bitext for two related low-resource African languages: Luhya and Swahili. For this work, we created a new parallel set of nearly 8000 Luhya-English sentences which allows a new zero-shot test of LASER and LaBSE. We find that LaBSE significantly outperforms LASER on both languages. Both LASER and LaBSE however perform poorly at zero-shot alignment on Luhya, achieving just 1.5% and 22.0% successful alignments respectively (P@1 score). We fine-tune the embeddings on a small set of parallel Luhya sentences and show significant gains, improving the LaBSE alignment accuracy to 53.3%. Further, restricting the dataset to sentence embedding pairs with cosine similarity above 0.7 yielded alignments with over 85% accuracy. Accepted to LoResMT 2022

Visit

arxiv.org

Tasks

embeddingsmachine translation

Languages

LuhyaSwahili

Tags

Computation and Language

Similar

Enhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word AlignmentLeveraging Closed-Access Multilingual Embedding for Automatic Sentence Alignment in Low Resource LanguagesEvaluating Sentence Alignment Methods in a Low-Resource Setting: An English-YorùBá Study CaseArtificial Code-Switching Training for Zero-Shot Cross-Lingual Sentence Embedding Alignment in Low-Resource African LanguagesVoice Conversion Can Improve ASR in Very Low-Resource SettingsOptimal Transport Distillation for Low-Resource Language Embedding Alignment

Enhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word Alignment

The field of cross-lingual sentence embeddings has recently experienced significant advancements, bu

Leveraging Closed-Access Multilingual Embedding for Automatic Sentence Alignment in Low Resource Languages

The importance of qualitative parallel data in machine translation has long been determined but it h

Evaluating Sentence Alignment Methods in a Low-Resource Setting: An English-YorùBá Study Case

Artificial Code-Switching Training for Zero-Shot Cross-Lingual Sentence Embedding Alignment in Low-Resource African Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Voice Conversion Can Improve ASR in Very Low-Resource Settings

Voice conversion (VC) could be used to improve speech recognition systems in low-resource languages

Optimal Transport Distillation for Low-Resource Language Embedding Alignment

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi