Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

GrEmLIn: A Repository of Green Baseline Embeddings for 87 Low-Resource Languages Injected with Multilingual Graph Knowledge

Domain:

natural language processing

Record type:

paperdatasetmodel
Creator:
GurKumOst
Host:avatar
Contextualized embeddings based on large language models (LLMs) are available for various languages, but their coverage is often limited for lower resourced languages. Using LLMs for such languages is often difficult due to a high computational cost; not only during training, but also during inference. Static word embeddings are much more resource-efficient ("green"), and thus still provide value, particularly for very low-resource languages. There is, however, a notable lack of comprehensive repositories with such embeddings for diverse languages. To address this gap, we present GrEmLIn, a centralized repository of green, static baseline embeddings for 87 mid- and low-resource languages. We compute GrEmLIn embeddings with a novel method that enhances GloVe embeddings by integrating multilingual graph knowledge, which makes our static embeddings competitive with LLM representations, while being parameter-free at inference time. Our experiments demonstrate that GrEmLIn embeddings outperform state-of-the-art contextualized embeddings from E5 on the task of lexical similarity. They remain competitive in extrinsic evaluation tasks like sentiment analysis and natural language inference, with average performance gaps of just 5-10\% or less compared to state-of-the-art models, given a sufficient vocabulary overlap with the target task, and underperform only on topic classification. Our code and embeddings are publicly available at huggingface.co. Long paper, accepted to NAACL 2025 Findings

Visit

arxiv.org

Tasks

embeddings

Tags

Computation and Language

Similar

Multilingual Word Embeddings for Low-Resource Languages using Anchors and a Chain of Related LanguagesmRAKL: Multilingual Retrieval-Augmented Knowledge Graph Construction for Low-Resourced LanguagesWALA: A Multilingual Resource Repository for West African LanguagesAfroXLMR-Comet: Multilingual Knowledge Distillation with Attention Matching for Low-Resource languagesMultilingual Knowledge Graphs and Low-Resource Languages: A ReviewMultilingual acoustic word embeddings for zero-resource languages

Multilingual Word Embeddings for Low-Resource Languages using Anchors and a Chain of Related Languages

Very low-resource languages, having only a few million tokens worth of data, are not well-supported

mRAKL: Multilingual Retrieval-Augmented Knowledge Graph Construction for Low-Resourced Languages

Knowledge Graphs represent real-world entities and the relationships between them. Multilingual Know

WALA: A Multilingual Resource Repository for West African Languages

The West African Language Archive (WALA) initiative has emerged from a number of concurrent projects, and aims to encourage local scholars to create high quality decentralised repositories documenting West African languages, and to make these repositories available

AfroXLMR-Comet: Multilingual Knowledge Distillation with Attention Matching for Low-Resource languages

Language model compression through knowledge distillation has emerged as a promising approach for de

Multilingual Knowledge Graphs and Low-Resource Languages: A Review

There is a lack of multilingual data to support applications in a large number of languages, especia

Multilingual acoustic word embeddings for zero-resource languages

This research addresses the challenge of developing speech applications for zero-resource languages