Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Unsupervised Cross-lingual Word Embedding Representation for English-isiZulu

Domaine:

natural language processing

Type de record:

modelpaper
Créateur:
AssAbbMabMarivate, Vukosi
Éditeur:
Und
Hôte:avatar
In this study, we investigate the effectiveness of using cross-lingual word embeddings for zero-shot transfer learning between a language with an abundant resource, English, and a language with limited resource, isiZulu. IsiZulu is a part of the South African Nguni language family, which is characterised by complex agglutinating morphology. We use VecMap, an open source tool, to obtain cross-lingual word embeddings. To perform an extrinsic evaluation of the effectiveness of the embeddings, we train a news classifier on labelled English data in order to categorise unlabelled isiZulu data using zero-shot transfer learning. In our study, we found our model to have a weighted average F1-score of 0.34. Our findings demonstrate that VecMap generates modular word embeddings in the cross-lingual space that have an impact on the downstream classifier used for zero-shot transfer learning.

Visit

doi.orgunderline.io

Tasks

embeddings

Languages

NgwoZulu

Tags

Natural Language ProcessingLanguage ModelsMachine LearningDeep Learning

Similaires

Unsupervised Cross-lingual Representation Learning at ScaleUnderstanding Linearity of Cross-Lingual Word Embedding MappingsEnhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word AlignmentPidginUNMT: Cross-lingual embedding between Pidgin and EnglishNamed Entity Recognition in Low-resource Languages using Cross-lingual distributional word representationImpact of Parallel Corpus Size on Cross-Lingual Word Embedding Quality in Low-Resource Languages

Unsupervised Cross-lingual Representation Learning at Scale

This paper shows that pretraining multilingual language models at scale leads to significant performance gains for a wide range of cross-lingual transfer tasks. We train a Transformer-based masked language model on one hundred languages, using more than two terabyt

Understanding Linearity of Cross-Lingual Word Embedding Mappings

The technique of Cross-Lingual Word Embedding (CLWE) plays a fundamental role in tackling Natural La

Enhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word Alignment

The field of cross-lingual sentence embeddings has recently experienced significant advancements, bu

PidginUNMT: Cross-lingual embedding between Pidgin and English

This repository contains the implementation of an Unsupervised NMT model from West African Pidgin (Creole) to English without using a single parallel sentence during training. Link to paper - https://arxiv.org/abs/1912.03444 (Accepted at NeurIPS 2019 Workshop on M

Named Entity Recognition in Low-resource Languages using Cross-lingual distributional word representation

Named Entity Recognition (NER) is a fundamental task in many NLP applications that seek to identify

Impact of Parallel Corpus Size on Cross-Lingual Word Embedding Quality in Low-Resource Languages

We propose a new approach for learning contextualised cross-lingual word embeddings based on a small