Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Enhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word Alignment

Domaine:

natural language processing

Type de record:

paper
Créateur:
MiaWu,ZhaWu,
Hôte:avatar
The field of cross-lingual sentence embeddings has recently experienced significant advancements, but research concerning low-resource languages has lagged due to the scarcity of parallel corpora. This paper shows that cross-lingual word representation in low-resource languages is notably under-aligned with that in high-resource languages in current models. To address this, we introduce a novel framework that explicitly aligns words between English and eight low-resource languages, utilizing off-the-shelf word alignment models. This framework incorporates three primary training objectives: aligned word prediction and word translation ranking, along with the widely used translation ranking. We evaluate our approach through experiments on the bitext retrieval task, which demonstrate substantial improvements on sentence embeddings in low-resource languages. In addition, the competitive performance of the proposed model across a broader range of tasks in high-resource languages underscores its practicality. NAACL 2024 findings

Visit

arxiv.org

Tasks

embeddings

Tags

Computation and Language

Similaires

Artificial Code-Switching Training for Zero-Shot Cross-Lingual Sentence Embedding Alignment in Low-Resource African LanguagesArtificial Code-Switching for Cross-Lingual Embedding Alignment in Low-Resource LanguagesCross-lingual NER Generalization via Embedding Alignment in Low-Resource LanguagesLeveraging Closed-Access Multilingual Embedding for Automatic Sentence Alignment in Low Resource LanguagesMultimodal Embedding Integration for Cross-Lingual NER in Low-Resource LanguagesImpact of Parallel Corpus Size on Cross-Lingual Word Embedding Quality in Low-Resource Languages

Artificial Code-Switching Training for Zero-Shot Cross-Lingual Sentence Embedding Alignment in Low-Resource African Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Artificial Code-Switching for Cross-Lingual Embedding Alignment in Low-Resource Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Cross-lingual NER Generalization via Embedding Alignment in Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Leveraging Closed-Access Multilingual Embedding for Automatic Sentence Alignment in Low Resource Languages

The importance of qualitative parallel data in machine translation has long been determined but it h

Multimodal Embedding Integration for Cross-Lingual NER in Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Impact of Parallel Corpus Size on Cross-Lingual Word Embedding Quality in Low-Resource Languages

We propose a new approach for learning contextualised cross-lingual word embeddings based on a small