Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Making a MIRACL: Multilingual Information Retrieval Across a Continuum of Languages

Domain:

natural language processing

Record type:

paperdataset
Creator:
ZhaThaOguKam
Host:avatar
MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual dataset we have built for the WSDM 2023 Cup challenge that focuses on ad hoc retrieval across 18 different languages, which collectively encompass over three billion native speakers around the world. These languages have diverse typologies, originate from many different language families, and are associated with varying amounts of available resources -- including what researchers typically characterize as high-resource as well as low-resource languages. Our dataset is designed to support the creation and evaluation of models for monolingual retrieval, where the queries and the corpora are in the same language. In total, we have gathered over 700k high-quality relevance judgments for around 77k queries over Wikipedia in these 18 languages, where all assessments have been performed by native speakers hired by our team. Our goal is to spur research that will improve retrieval across a continuum of languages, thus enhancing information access capabilities for diverse populations around the world, particularly those that have been traditionally underserved. This overview paper describes the dataset and baselines that we share with the community. The MIRACL website is live at miracl.ai.

Visit

arxiv.org

Tasks

information retrieval

Tags

Information RetrievalComputation and Language

Similar

Comparative Analysis of Hybrid Batch and Standard Multilingual Dense Retrieval on MIRACL for Low-Resource LanguagesMultilingual Information Retrieval with a Monolingual Knowledge BaseMIRACL Retrieval Accuracy via Combined Monolingual, Cross-Lingual, and Multilingual Data Augmentation for Low-Resource LanguagesDETECTING CYBERBULLYING ACROSS NIGERIAN LANGUAGES: A MULTILINGUAL SYSTEMLeveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense RetrievalInformation Retrieval in African Languages

Comparative Analysis of Hybrid Batch and Standard Multilingual Dense Retrieval on MIRACL for Low-Resource Languages

Information retrieval across different languages is an increasingly important challenge in natural l

Multilingual Information Retrieval with a Monolingual Knowledge Base

Multilingual information retrieval has emerged as powerful tools for expanding knowledge sharing acr

MIRACL Retrieval Accuracy via Combined Monolingual, Cross-Lingual, and Multilingual Data Augmentation for Low-Resource Languages

Information retrieval across different languages is an increasingly important challenge in natural l

DETECTING CYBERBULLYING ACROSS NIGERIAN LANGUAGES: A MULTILINGUAL SYSTEM

ABSTRACT

As the digital world evolves, the

Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval

There has been limited success for dense retrieval models in multilingual retrieval, due to uneven a

Information Retrieval in African Languages

Developing Information Retrieval (IR) tools and techniques in African languages suffers from the dua