Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

African Next Voices: Pilot Data Collection in Kenya

Domain:

agricultureeducationsocioeconomic

Record type:

dataset
African Next Voices: Pilot Data Collection in Kenya is part of a larger initiative to support African language speech technology. This project, funded by the Gates Foundation, is led by the KenCorpus Consortium, a coalition of Kenyan universities and research centers. It includes scripted and unscripted speech across multiple domains and five languages, collected through ethical, community-led processes.

Visit

huggingface.coDholuoKikuyuSomaliMaasaiKalenjin"

Languages

GikuyuKalenjinKipsigisMaasaiSomali

Tags

African Next VoicesANVASRKenya

Licenses

CC BY 4.0

Similar

African Next Voices: Kenya and Tanzaniaza-african-next-voicesza-african-next-voicesMali African Next VoicesRwanda African Next VoicesAfrican Next Voices: Ethiopia

African Next Voices: Kenya and Tanzania

Afrivoice_Swahili is an open-source image-prompt speech corpus for Swahili ASR development. It contains ~3,000 hours of audio spanning five domains: Agriculture, Education, Finance, Government, and Health.

za-african-next-voices

Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 S

za-african-next-voices

Note: This dataset is a compressed version of za-african-next-voices. It was compressed to .opus for

Mali African Next Voices

The AfVoices dataset is the largest open corpus of spontaneous Bambara speech at its release in late 2025. It contains 423 hours of segmented audio and 612 hours of original raw recordings collected across southern Mali. Speech was recorded in natural, conversation

Rwanda African Next Voices

The dataset was created by Digital Umuganda and made possible through funding from the Gates Foundation. The data spans five high-impact domains — Health, Government, Financial Services, Education, and Agriculture — to support robust ASR model development in both c

African Next Voices: Ethiopia

AfriVoice Ethiopia is an open-source speech corpus for ASR development covering five Ethiopian languages: Amharic, Afaan Oromo, Sidama, Wolaytta, and Tigrinya.