Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

AfriVoices-KE: A Multilingual Speech Dataset for Kenyan Languages

Domain:

natural language processing

Record type:

paperdataset
Creator:
Wanzare, LilianAmoMaiOdh
Host:avatar
AfriVoices-KE is a large-scale multilingual speech dataset comprising approximately 3,000 hours of audio across five Kenyan languages: Dholuo, Kikuyu, Kalenjin, Maasai, and Somali. The dataset includes 750 hours of scripted speech and 2,250 hours of spontaneous speech, collected from 4,777 native speakers across diverse regions and demographics. This work addresses the critical underrepresentation of African languages in speech technology by providing a high-quality, linguistically diverse resource. Data collection followed a dual methodology: scripted recordings drew from compiled text corpora, translations, and domain-specific generated sentences spanning eleven domains relevant to the Kenyan context, while unscripted speech was elicited through textual and image prompts to capture natural linguistic variation and dialectal nuances. A customized mobile application enabled contributors to record using smartphones. Quality assurance operated at multiple layers, encompassing automated signal-to-noise ratio validation prior to recording and human review for content accuracy. Though the project encountered challenges common to low-resource settings, including unreliable infrastructure, device compatibility issues, and community trust barriers, these were mitigated through local mobilizers, stakeholder partnerships, and adaptive training protocols. AfriVoices-KE provides a foundational resource for developing inclusive automatic speech recognition and text-to-speech systems, while advancing the digital preservation of Kenya's linguistic heritage. 10 pages, 5 figures, 3 tables

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Languages

DholuoGikuyuKalenjinKipsigisMaasaiSomali

Tags

Computation and Language

Similar

UGSpeechData: A Multilingual Speech Dataset of Ghanaian Languages UGSpeechData: A Multilingual Speech Dataset of Ghanaian LanguagesAfrican Voices: Multilingual Speech Dataset for Low-Resource African LanguagesKenPos: Kenyan Languages Part of Speech Tagged datasetFikira Dataset | A Multilingual Reasoning Dataset for African LanguagesA multilingual dataset for offensive language and hate speech detection for hausa, yoruba and igbo languagesZambezi Voice: A Multilingual Speech Corpus for Zambian Languages

UGSpeechData: A Multilingual Speech Dataset of Ghanaian Languages UGSpeechData: A Multilingual Speech Dataset of Ghanaian Languages

The UGSpeechData is a collection of audio speech data of Akan, Ewe, Dagaare, Dagbani, and Ikposo. Th

African Voices: Multilingual Speech Dataset for Low-Resource African Languages

A large-scale multilingual speech dataset developed by Data Science Nigeria. Contains more than 3,000 hours of transcribed audio across four Nigerian languages: Hausa, Igbo, Nigerian Pidgin, and Yorùbá. The dataset supports Automatic Speech Recognition (ASR) and sp

KenPos: Kenyan Languages Part of Speech Tagged dataset

This project developed a Part of Speech (POS) Tagged dataset of 2 languages in Kenya: Dholuo and 3 Luhya dialects (Lumarachi, Lulogooli, and Lubukusi). The project tagged approximately 143,000 words, which includes about 50,000 words for Dholuo, 27,900 words for Lu

Fikira Dataset | A Multilingual Reasoning Dataset for African Languages

Fikira (Swahili for "thinking/reasoning") is a multilingual reasoning dataset for African languages, developed by Vambo AI. This dataset contains 50,000 reasoning examples across 10 African languages, synthetically generated as part of ongoing experiments at Vambo

A multilingual dataset for offensive language and hate speech detection for hausa, yoruba and igbo languages

The proliferation of online offensive language necessitates the development of effective detection m

Zambezi Voice: A Multilingual Speech Corpus for Zambian Languages

This work introduces Zambezi Voice, an open-source multilingual speech resource for Zambian languages. It contains two collections of datasets: unlabelled audio recordings of radio news and talk shows programs (160 hours) and labelled data (over 80 hours) consistin