Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

AfriMTEB and AfriE5: Benchmarking and Adapting Text Embedding Models for African Languages

Domain:

natural language processing

Record type:

paperdatasetmodelsoftware
Creator:
UemZhaAde
Host:avatar
Text embeddings are an essential building component of several NLP tasks such as retrieval-augmented generation which is crucial for preventing hallucinations in LLMs. Despite the recent release of massively multilingual MTEB (MMTEB), African languages remain underrepresented, with existing tasks often repurposed from translation benchmarks such as FLORES clustering or SIB-200. In this paper, we introduce AfriMTEB -- a regional expansion of MMTEB covering 59 languages, 14 tasks, and 38 datasets, including six newly added datasets. Unlike many MMTEB datasets that include fewer than five languages, the new additions span 14 to 56 African languages and introduce entirely new tasks, such as hate speech detection, intent detection, and emotion classification, which were not previously covered. Complementing this, we present AfriE5, an adaptation of the instruction-tuned mE5 model to African languages through cross-lingual contrastive distillation. Our evaluation shows that AfriE5 achieves state-of-the-art performance, outperforming strong baselines such as Gemini-Embeddings and mE5. Accepted to EACL 2026 (main conference)

Visit

arxiv.org

Tasks

embeddingsemotion identificationhate speech detectiontext classification

Tags

Computation and Language

Similar

Benchmarking Text Embedding Models for South African LanguagesNGLUEni: Benchmarking and Adapting Pretrained Language Models for Nguni LanguagesAfroXLMR-Social: Adapting Pre-trained Language Models for African Languages Social Media TextBenchmarking Automatic Speech Recognition Models for African LanguagesLugha-Llama: Adapting Large Language Models for African LanguagesOptimized Text Embedding Models and Benchmarks for Amharic Passage Retrieval

Benchmarking Text Embedding Models for South African Languages

NGLUEni: Benchmarking and Adapting Pretrained Language Models for Nguni Languages

AfroXLMR-Social: Adapting Pre-trained Language Models for African Languages Social Media Text

Language models built from various sources are the foundation of today's NLP progress. However, for

Benchmarking Automatic Speech Recognition Models for African Languages

Automatic speech recognition (ASR) for African languages remains constrained by limited labeled data

Lugha-Llama: Adapting Large Language Models for African Languages

Large language models (LLMs) have achieved impressive results in a wide range of natural language ap

Optimized Text Embedding Models and Benchmarks for Amharic Passage Retrieval

Neural retrieval methods using transformer-based pre-trained language models have advanced multiling