Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Combining Pretrained High-Resource Embeddings and Subword Representations for Low-Resource Languages

Domaine:

natural language processing

Type de record:

paper
The contrast between the need for large amounts of data for current Natural Language Processing (NLP) techniques, and the lack thereof, is accentuated in the case of African languages, most of which are considered low-resource. To help circumvent this issue, we explore techniques exploiting the qualities of morphologically rich languages (MRLs), while leveraging pretrained word vectors in well-resourced languages. In our exploration, we show that a meta-embedding approach combining both pretrained and morphologically-informed word embeddings performs best in the downstream task of Xhosa-English translation.

Visit

arxiv.org

Tasks

embeddingsmachine translation

Languages

Xhosa

Tags

africanlp1

Similaires

Isomorphic Cross-lingual Embeddings for Low-Resource LanguagesBitext Mining Using Distilled Sentence Representations for Low-Resource LanguagesScaling Pretrained Models and Intermediate-Task Training for Low-Resource Languages in XTREMEMorphological Segmentation to Improve Crosslingual Word Embeddings for Low Resource LanguagesOn the Importance of Subword Information for Morphological Tasks in Truly Low-Resource LanguagesMorphology-aware Subword Segmentation for Zero-shot Cross-lingual Transfer in Low-resource Languages

Isomorphic Cross-lingual Embeddings for Low-Resource Languages

Cross-Lingual Word Embeddings (CLWEs) are a key component to transfer linguistic information learnt

Bitext Mining Using Distilled Sentence Representations for Low-Resource Languages

Scaling multilingual representation learning beyond the hundred most frequent languages is challengi

Scaling Pretrained Models and Intermediate-Task Training for Low-Resource Languages in XTREME

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Morphological Segmentation to Improve Crosslingual Word Embeddings for Low Resource Languages

Crosslingual word embeddings developed from multiple parallel corpora help in understanding the rela

On the Importance of Subword Information for Morphological Tasks in Truly Low-Resource Languages

Recent work has validated the importance of subword information for word representation learning. Si

Morphology-aware Subword Segmentation for Zero-shot Cross-lingual Transfer in Low-resource Languages

Multilingual modelling can improve machine translation for low-resource languages, partly through sh