African Next Voices: Pilot Data Collection in Kenya is part of a larger initiative to support African language speech technology. This project, funded by the Gates Foundation, is led by the KenCorpus Consortium, a coalition of Kenyan universities and research cente
Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 S
Note: This dataset is a compressed version of za-african-next-voices. It was compressed to .opus for
The AfVoices dataset is the largest open corpus of spontaneous Bambara speech at its release in late 2025. It contains 423 hours of segmented audio and 612 hours of original raw recordings collected across southern Mali. Speech was recorded in natural, conversation
The dataset was created by Digital Umuganda and made possible through funding from the Gates Foundation. The data spans five high-impact domains — Health, Government, Financial Services, Education, and Agriculture — to support robust ASR model development in both c
AfriVoice Ethiopia is an open-source speech corpus for ASR development covering five Ethiopian languages: Amharic, Afaan Oromo, Sidama, Wolaytta, and Tigrinya.