Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Using Subword-Embeddings for Bilingual Lexicon Induction in Bantu Languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
AssAkbJoh
Éditeur:
Und
Hôte:avatar
Bilingual Lexicon Induction (BLI) is a valuable tool in machine translation and cross-lingual transfer learning, but it remains challenging for agglutinative and low-resource languages. In this work, we investigate the use of weighted sub-word embeddings in BLI for agglutinative languages. We further evaluate a graph-matching and Procrustes-based BLI approach on two Bantu languages, assessing its effectiveness in a previously underexplored language family. Our results for Swahili with an average P@1 score of $51.84$% for a $3000$ word dictionary demonstrate the success of the approach for Bantu languages. Weighted sub-word embeddings perform competitively on Swahili and outperform word embeddings in our experiments with Zulu.

Visit

doi.orgunderline.io

Tasks

embeddingsmachine translation

Languages

Swahili

Tags

Computational LinguisticsNatural Language ProcessingArtificial Intelligence

Similaires

Combining Pretrained High-Resource Embeddings and Subword Representations for Low-Resource LanguagesHierarchical Multi Task Learning with Subword Contextual Embeddings for Languages with Rich MorphologyResource-Lean Lexicon Induction for German DialectsExpanding Swahili Lexicon by Means of Bantu LanguagesUnsupervised POS Induction with Word EmbeddingsScaling Bilingual Lexicon Size for Robust Zero-Shot Cross-Lingual Retrieval in Low-Resource Languages

Combining Pretrained High-Resource Embeddings and Subword Representations for Low-Resource Languages

The contrast between the need for large amounts of data for current Natural Language Processing (NLP) techniques, and the lack thereof, is accentuated in the case of African languages, most of which are considered low-resource. To help circumvent this issue, we exp

Hierarchical Multi Task Learning with Subword Contextual Embeddings for Languages with Rich Morphology

Morphological information is important for many sequence labeling tasks in Natural Language Processi

Resource-Lean Lexicon Induction for German Dialects

Automatic induction of high-quality dictionaries is essential for building lexical resources, yet lo

Expanding Swahili Lexicon by Means of Bantu Languages

Unsupervised POS Induction with Word Embeddings

Unsupervised word embeddings have been shown to be valuable as features in supervised learning probl

Scaling Bilingual Lexicon Size for Robust Zero-Shot Cross-Lingual Retrieval in Low-Resource Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to