Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Impact of Parallel Corpus Size on Cross-Lingual Word Embedding Quality in Low-Resource Languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
Ass
Éditeur:
Zenodo
Hôte:avatar
We propose a new approach for learning contextualised cross-lingual word embeddings based on a small parallel corpus (e.g. a few hundred sentence pairs). Our method obtains word embeddings via an LSTM encoder-decoder model that simultaneously translates and reconstructs an input sentence. Through sharing model parameters among different languages, our model jointly trains the word embeddings in a common cross-lingual space. We also propose to combine word and subword embeddings to make use of orthographic similarities across different languages. We base our experiments on real-world data from Research goal: What is the impact of varying the size of parallel corpora (e.g., 100 vs. 1,000 sentence pairs) on the quality of contextualized cross-lingual word embeddings when evaluated on the XNLI benchmark for low-resource languages? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.3/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 8.3/10.

Visit

doi.org

Tasks

embeddings

Tags

impactvaryingsizeparallelcorporasentencepairsquality

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Impact of Pre-training Corpus Size on Cross-lingual Entity Recognition in Low-resource LanguagesEnhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word AlignmentCross-lingual NER Generalization via Embedding Alignment in Low-Resource LanguagesMultimodal Embedding Integration for Cross-Lingual NER in Low-Resource LanguagesLearning Contextualised Cross-lingual Word Embeddings and Alignments for Extremely Low-Resource Languages Using Parallel CorporaImpact of Intermediate Task Training Data Size on Zero-Shot Cross-Lingual Transfer Accuracy in Low-Resource Languages

Impact of Pre-training Corpus Size on Cross-lingual Entity Recognition in Low-resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Enhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word Alignment

The field of cross-lingual sentence embeddings has recently experienced significant advancements, bu

Cross-lingual NER Generalization via Embedding Alignment in Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Multimodal Embedding Integration for Cross-Lingual NER in Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Learning Contextualised Cross-lingual Word Embeddings and Alignments for Extremely Low-Resource Languages Using Parallel Corpora

We propose a new approach for learning contextualised cross-lingual word embeddings based on a small

Impact of Intermediate Task Training Data Size on Zero-Shot Cross-Lingual Transfer Accuracy in Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni