Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

rasyosef/Amharic-Passage-Retrieval-Dataset-V2

Domain:

natural language processing

Record type:

dataset
Creator:
ras
Host:
This dataset is the official benchmark for the paper "The Multilingual Curse at the Retrieval Layer: Evidence from Amharic". It provides a fixed 90/10 train–test split consisting of 68,000 query–passage pairs, specifically designed for evaluating and training dense, late-interaction, learned sparse, and cross-encoder retrieval models for the Amharic language. GitHub Repository: rasyosef/amharic-neural-ir

Visit

huggingface.co

Tasks

information retrieval

Languages

Amharic

Licenses

cc-by-nc-sa-4.0

Similar

rasyosef/Amharic-Passage-Retrieval-Dataset-V2-With-Negativesrasyosef/amharic-passage-retrieval-datasetrasyosef/amharic-passage-retrieval-dataset-with-negativesyosefw/amharic-passage-retrieval-dataset-v2-with-negativesDesalegnn/amharic-passage-retrieval-datasetDesalegnn/new-amharic-passage-retrieval-dataset

rasyosef/Amharic-Passage-Retrieval-Dataset-V2-With-Negatives

This dataset can be used directly with Sentence Transformers to train Amharic Embedding and Rerankin

rasyosef/amharic-passage-retrieval-dataset

This dataset is a version of amharic-news-category-classification that has been filtered, deduplicat

rasyosef/amharic-passage-retrieval-dataset-with-negatives

This dataset is a version of amharic-news-category-classification that has been filtered, deduplicat

yosefw/amharic-passage-retrieval-dataset-v2-with-negatives

Desalegnn/amharic-passage-retrieval-dataset

Desalegnn/new-amharic-passage-retrieval-dataset

This dataset is generated from AMQA-style question–context pairs, converted to match the schema of D