Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

KenLumachiQuAD – A Question Answering Dataset for Kenyan Luhya Lumarachi Language for Machine Learning

Domain:

natural language processing

Record type:

dataset
Creator:
BarLawEva
Publisher:
Eas
Host:
Question-answering (QA) datasets play a crucial role in testing and training machine learning models, from which we can develop practical end-user applications, such as internet search, dialogue systems, and chatbots. There are various methods for creating QA datasets, including the use of transfer learning, employing translations, creation from synthetic data, and the use of human annotators, depending on the availability of reference datasets. QA datasets for low-resource languages are few and tend to be of small sizes due to the costs associated with human annotation, which is the easiest available method for such low-resource languages in the absence of other methods that need some existing reference data. We make our contribution in the provision of QA resources for low-resource languages of Kenya by developing a human-annotated QA dataset, called KenLumachiQuAD.  This is a QA dataset for the Kenyan low-resource language of Luhya, specifically the Lumarachi dialect. KenLumachiQuAD is a dataset of 1,000 QA pairs that is human-annotated from available public domain texts for the Luhya Lumarachi language. We evaluated the dataset on its applicability to the machine learning task of QA using a semantic network modelling method, based on a small sample of the dataset, and achieved a result of 76% exact match. The research, therefore, provides a QA dataset for machine learning and contributes to the resourcing of low-resource languages. Researchers can still add more QA to this dataset from texts that are yet to be annotated to further augment this dataset. Finally, other researchers can apply our experience in developing other QA datasets for low-resource languages

Visit

doi.org

Tasks

question answering

Languages

Luhya

Licenses

http://creativecommons.org/licenses/by/4.0

Similar

KenLumachiQuAD - A QA dataset for Kenyan Luhya Lumarachi dialect A Swahili Question-Answering Dataset for Machine Reading Comprehension in HorticultureKenSwQuAD – A Question Answering Dataset for Swahili Low Resource LanguageSwahiliVQA: A Dataset for Visual Question Answering in Swahili Language

KenLumachiQuAD - A QA dataset for Kenyan Luhya Lumarachi dialect

KenLumachiQuAD is a result of a project that annotated a total of 1,000 QA pairs based on 137 texts

A Swahili Question-Answering Dataset for Machine Reading Comprehension in Horticulture

KenSwQuAD – A Question Answering Dataset for Swahili Low Resource Language

This research developed a Kencorpus Swahili Question Answering Dataset KenSwQuAD from raw data of Swahili language, which is a low resource language predominantly spoken in Eastern African and also has speakers in other parts of the world. Question Answering datase

SwahiliVQA: A Dataset for Visual Question Answering in Swahili Language