Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

SwaQuAD-24: QA Benchmark Dataset in Swahili

Domain:

natural language processing

Record type:

paperdataset
Creator:
Kon
Host:avatar
This paper proposes the creation of a Swahili Question Answering (QA) benchmark dataset, aimed at addressing the underrepresentation of Swahili in natural language processing (NLP). Drawing from established benchmarks like SQuAD, GLUE, KenSwQuAD, and KLUE, the dataset will focus on providing high-quality, annotated question-answer pairs that capture the linguistic diversity and complexity of Swahili. The dataset is designed to support a variety of applications, including machine translation, information retrieval, and social services like healthcare chatbots. Ethical considerations, such as data privacy, bias mitigation, and inclusivity, are central to the dataset development. Additionally, the paper outlines future expansion plans to include domain-specific content, multimodal integration, and broader crowdsourcing efforts. The Swahili QA dataset aims to foster technological innovation in East Africa and provide an essential resource for NLP research and applications in low-resource languages.

Visit

arxiv.org

Tasks

question answering

Languages

Swahili

Tags

Computation and LanguageArtificial Intelligence

Similar

KenSwQuAD: Swahili QA DatasetAfriMed-QA: A Pan-African, Multi-Specialty, Medical Question-Answering Benchmark Datasetmgudle/bert-swahili-qaAmharic STI QA DatasetJasperV13/MoroccanHistory-QA-Darija-DatasetEthiopian Family Code QA Dataset

KenSwQuAD: Swahili QA Dataset

Question answering systems, NLP research Notes / challenges: Limited to Swahili; dataset size may b

AfriMed-QA: A Pan-African, Multi-Specialty, Medical Question-Answering Benchmark Dataset

Recent advancements in large language model(LLM) performance on medical multiple choice question (MC

mgudle/bert-swahili-qa

Amharic STI QA Dataset

JasperV13/MoroccanHistory-QA-Darija-Dataset

This dataset is a translated subset from KBayoud/MoroccanHistory-QA-Dataset.The translation from Eng

Ethiopian Family Code QA Dataset

This dataset contains collection of question-and-answer pairs that have been collected in two ways.