Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

DistillEmb: Distilling word embeddings via contrastive learning

Domain:

natural language processing

Record type:

paper
Creator:
AssMerWu,
Publisher:
Und
Host:avatar
Word embeddings powered the early days of neural network-based NLP research. Their effectiveness in small data regimes makes them still relevant in low-resource environments. However, they are limited in two critical ways: linearly increasing memory requirement based on the number of tokens and out-of-vocabulary token handling. In this work, we present a distillation technique of word embeddings into a CNN network using contrastive learning. This method allows embeddings to be regressed given the characters of a token. Low resource languages are the primary beneficiary of these distilled embeddings and hence, we show the effectiveness of such a model on Amharic, Semitic languages that is spoken in Ethiopia.

Visit

doi.orgunderline.io

Tasks

embeddings

Languages

Amharic

Tags

Natural Language ProcessingLanguage Models

Similar

KinyaEmbed: Contrastive Sentence Embeddings for Kinyarwanda via Multi-Stage Curriculum TrainingLearning Multilingual Word Embeddings Using Image-Text DataBitext Mining for Low-Resource Languages via Contrastive LearningCompass-Embedding v4: Robust Contrastive Learning for Multilingual E-commerce EmbeddingsisiZulu Word EmbeddingsImproving Cross-Lingual NER Robustness via Contrastive Learning in XTREME-R

KinyaEmbed: Contrastive Sentence Embeddings for Kinyarwanda via Multi-Stage Curriculum Training

We present KinyaEmbed, the first dedicated sentence embedding model for Kinyarwanda, a morphological

Learning Multilingual Word Embeddings Using Image-Text Data

There has been significant interest recently in learning multilingual word embeddings -- in which se

Bitext Mining for Low-Resource Languages via Contrastive Learning

Mining high-quality bitexts for low-resource languages is challenging. This paper shows that sentenc

Compass-Embedding v4: Robust Contrastive Learning for Multilingual E-commerce Embeddings

As global e-commerce rapidly expands into emerging markets, the lack of high-quality semantic repres

isiZulu Word Embeddings

Improving Cross-Lingual NER Robustness via Contrastive Learning in XTREME-R

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident