Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

AmQA: Amharic Question Answering Dataset

Domain:

natural language processing

Record type:

paperdataset
Creator:
AbeUsbAss
Host:avatar
Question Answering (QA) returns concise answers or answer lists from natural language text given a context document. Many resources go into curating QA datasets to advance robust models' development. There is a surge of QA datasets for languages like English, however, this is not true for Amharic. Amharic, the official language of Ethiopia, is the second most spoken Semitic language in the world. There is no published or publicly available Amharic QA dataset. Hence, to foster the research in Amharic QA, we present the first Amharic QA (AmQA) dataset. We crowdsourced 2628 question-answer pairs over 378 Wikipedia articles. Additionally, we run an XLMR Large-based baseline model to spark open-domain QA research interest. The best-performing baseline achieves an F-score of 69.58 and 71.74 in reader-retriever QA and reading comprehension settings respectively.

Visit

arxiv.org

Tasks

question answering

Languages

Amharic

Tags

Computation and LanguageArtificial IntelligenceInformation Retrieval

Similar

Amharic health question answering datasetLow Resource Question Answering: An Amharic Benchmarking DatasetzeharaY/Amharic-question-answeringAVQANet: Amharic Visual Question Answering Dataset for Ethiopian Museum ArtifactsTigrinya Question-Answering Dataset (TiQuAD)Swahili Visual Question Answering Dataset

Amharic health question answering dataset

The **AmHQA** dataset is an Amharic health question answering corpus curated to support research in

Low Resource Question Answering: An Amharic Benchmarking Dataset

zeharaY/Amharic-question-answering

# Amharic-question-answering

AVQANet: Amharic Visual Question Answering Dataset for Ethiopian Museum Artifacts

This dataset was developed for research on Amharic Visual Question Answering (AVQA) systems for Ethi

Tigrinya Question-Answering Dataset (TiQuAD)

TiQuAD is a human-annotated question-answering dataset for the Tigrinya language. The dataset contai

Swahili Visual Question Answering Dataset

Swahili Language VQA Dataset