Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

mteb/MIRACLRetrieval

Domain:

natural language processing

Record type:

dataset
Creator:
mteb
Host:
MIRACLRetrieval An MTEB dataset Massive Text Embedding Benchmark MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual retrieval dataset that focuses on search across 18 different languages. Task category t2t Domains Encyclopaedic, Written Reference miracl.ai You can evaluate an embedding model on this dataset using the following code: import mteb

Visit

huggingface.co

Tasks

information retrieval

Languages

SwahiliYoruba

Tags

mtebtextlarge datasets from Lanfrica Insights

Licenses

cc-by-sa-4.0

Similar

mteb/InjongoIntentmteb/SiswatiNewsClassificationmteb/AfriSentiClassificationmteb/AfriHateClassificationmteb/siswati_newsmteb/SemRel24STS

mteb/InjongoIntent

InjongoIntent An MTEB dataset Massive Text Embedding Benchmark Multicultural intent-classification

mteb/SiswatiNewsClassification

SiswatiNewsClassification An MTEB dataset Massive Text Embedding Benchmark Siswati News Classificat

mteb/AfriSentiClassification

AfriSentiClassification An MTEB dataset Massive Text Embedding Benchmark AfriSenti is the largest s

mteb/AfriHateClassification

AfriHateClassification An MTEB dataset Massive Text Embedding Benchmark AfriHate is a multilingual

mteb/siswati_news

mteb/SemRel24STS

SemRel24STS An MTEB dataset Massive Text Embedding Benchmark SemRel2024 is a collection of Semantic