Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Multilingual Word Embeddings for Low-Resource Languages using Anchors and a Chain of Related Languages

Domain:

natural language processing

Record type:

paper
Creator:
HanSevRalFra
Host:avatar
Very low-resource languages, having only a few million tokens worth of data, are not well-supported by multilingual NLP approaches due to poor quality cross-lingual word representations. Recent work showed that good cross-lingual performance can be achieved if a source language is related to the low-resource target language. However, not all language pairs are related. In this paper, we propose to build multilingual word embeddings (MWEs) via a novel language chain-based approach, that incorporates intermediate related languages to bridge the gap between the distant source and target. We build MWEs one language at a time by starting from the resource rich source and sequentially adding each language in the chain till we reach the target. We extend a semi-joint bilingual approach to multiple languages in order to eliminate the main weakness of previous works, i.e., independently trained monolingual embeddings, by anchoring the target language around the multilingual space. We evaluate our method on bilingual lexicon induction for 4 language families, involving 4 very low-resource (<5M tokens) and 4 moderately low-resource (<50M) target languages, showing improved performance in both categories. Additionally, our analysis reveals the importance of good quality embeddings for intermediate languages as well as the importance of leveraging anchor points from all languages in the multilingual space. Accepted at the MRL 2023 workshop

Visit

arxiv.org

Tasks

embeddings

Tags

Computation and Language

Similar

Multilingual acoustic word embeddings for zero-resource languagesMorphological Segmentation to Improve Crosslingual Word Embeddings for Low Resource LanguagesLearning Contextualised Cross-lingual Word Embeddings and Alignments for Extremely Low-Resource Languages Using Parallel CorporaMultilingual transfer of acoustic word embeddings improves when training on languages related to the target zero-resource languageMultilingual NLP for Low-Resource Languages Using Transfer LearningA Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages

Multilingual acoustic word embeddings for zero-resource languages

This research addresses the challenge of developing speech applications for zero-resource languages

Morphological Segmentation to Improve Crosslingual Word Embeddings for Low Resource Languages

Crosslingual word embeddings developed from multiple parallel corpora help in understanding the rela

Learning Contextualised Cross-lingual Word Embeddings and Alignments for Extremely Low-Resource Languages Using Parallel Corpora

We propose a new approach for learning contextualised cross-lingual word embeddings based on a small

Multilingual transfer of acoustic word embeddings improves when training on languages related to the target zero-resource language

Acoustic word embedding models map variable duration speech segments to fixed dimensional vectors, enabling efficient speech search and discovery. Previous work explored how embeddings can be obtained in zero-resource settings where no labelled data is available in

Multilingual NLP for Low-Resource Languages Using Transfer Learning

Abstract: Despite the emergence of large-scale multilingual pre-trained models like mBERT, XLM-RoBER

A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages

We use the multilingual OSCAR corpus, extracted from Common Crawl via language classification, filtering and cleaning, to train monolingual contextualized word embeddings (ELMo) for five mid-resource languages. We then compare the performance of OSCAR-based and Wik