Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Very Low Resource Sentence Alignment: Luhya and Swahili

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
ChiBas
Hôte:avatar
Language-agnostic sentence embeddings generated by pre-trained models such as LASER and LaBSE are attractive options for mining large datasets to produce parallel corpora for low-resource machine translation. We test LASER and LaBSE in extracting bitext for two related low-resource African languages: Luhya and Swahili. For this work, we created a new parallel set of nearly 8000 Luhya-English sentences which allows a new zero-shot test of LASER and LaBSE. We find that LaBSE significantly outperforms LASER on both languages. Both LASER and LaBSE however perform poorly at zero-shot alignment on Luhya, achieving just 1.5% and 22.0% successful alignments respectively (P@1 score). We fine-tune the embeddings on a small set of parallel Luhya sentences and show significant gains, improving the LaBSE alignment accuracy to 53.3%. Further, restricting the dataset to sentence embedding pairs with cosine similarity above 0.7 yielded alignments with over 85% accuracy. Accepted to LoResMT 2022

Visit

arxiv.org

Tasks

embeddingsmachine translation

Languages

LuhyaSwahili

Tags

Computation and Language

Similaires

Enhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word AlignmentLeveraging Closed-Access Multilingual Embedding for Automatic Sentence Alignment in Low Resource LanguagesEvaluating Sentence Alignment Methods in a Low-Resource Setting: An English-YorùBá Study CaseArtificial Code-Switching Training for Zero-Shot Cross-Lingual Sentence Embedding Alignment in Low-Resource African LanguagesVoice Conversion Can Improve ASR in Very Low-Resource SettingsOptimal Transport Distillation for Low-Resource Language Embedding Alignment

Enhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word Alignment

The field of cross-lingual sentence embeddings has recently experienced significant advancements, bu

Leveraging Closed-Access Multilingual Embedding for Automatic Sentence Alignment in Low Resource Languages

The importance of qualitative parallel data in machine translation has long been determined but it h

Evaluating Sentence Alignment Methods in a Low-Resource Setting: An English-YorùBá Study Case

Artificial Code-Switching Training for Zero-Shot Cross-Lingual Sentence Embedding Alignment in Low-Resource African Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Voice Conversion Can Improve ASR in Very Low-Resource Settings

Voice conversion (VC) could be used to improve speech recognition systems in low-resource languages

Optimal Transport Distillation for Low-Resource Language Embedding Alignment

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi