Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

AfriCLIRMatrix: Enabling Cross-Lingual Information Retrieval for African Languages

Domain:

natural language processing

Record type:

paper

Language diversity in NLP is critical in enabling the development of tools for a wide range of users.However, there are limited resources for building such tools for many languages, particularly those spoken in Africa.For search, most existing datasets feature few or no African languages, directly impacting researchers’ ability to build and improve information access capabilities in those languages.Motivated by this, we created AfriCLIRMatrix, a test collection for cross-lingual information retrieval research in 15 diverse African languages.In total, our dataset contains 6 million queries in English and 23 million relevance judgments automatically mined from Wikipedia inter-language links, covering many more African languages than any existing information retrieval test collection.In addition, we release BM25, dense retrieval, and sparse–dense hybrid baselines to provide a starting point for the development of future systems. We hope that these efforts can spur additional work in search for African languages.AfriCLIRMatrix can be downloaded at AfriCLIRMatrix.

Visit

aclanthology.orguwaterloo.ca

Connected records

dataset

Tasks

question answeringinformation retrieval

Languages

AfrikaansAkanAmharicArabic, Egyptian SpokenArabic, Moroccan SpokenBwamu, CwiDinka, SoutheasternHausaIgboShona+7

Tags

Cross-Lingual Information Retrieval

Similar

AfriQA: Cross-lingual Open-Retrieval Question Answering for African LanguagesData Efficient Dense Cross-Lingual Information RetrievalCross-Lingual Retrieval Augmented Prompt for Low-Resource LanguagesImproving Cross-lingual Information Retrieval on Low-Resource Languages via Optimal Transport DistillationInformation Retrieval in African LanguagesQuery Augmentation for Cross-Lingual Dense Retrieval in Low-Resource Languages

AfriQA: Cross-lingual Open-Retrieval Question Answering for African Languages

African languages have far less in-language content available digitally, making it challenging for q

Data Efficient Dense Cross-Lingual Information Retrieval

Cross-Lingual Information Retrieval (CIR) remains challenging due to limited annotated data and ling

Cross-Lingual Retrieval Augmented Prompt for Low-Resource Languages

Multilingual Pretrained Language Models (MPLMs) have shown their strong multilinguality in recent em

Improving Cross-lingual Information Retrieval on Low-Resource Languages via Optimal Transport Distillation

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

Information Retrieval in African Languages

Developing Information Retrieval (IR) tools and techniques in African languages suffers from the dua

Query Augmentation for Cross-Lingual Dense Retrieval in Low-Resource Languages

Effective cross-lingual dense retrieval methods that rely on multilingual pre-trained language model