Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 S
Swivuriso is a 3000-hour multilingual speech dataset developed as part of the African Next Voices project, to support the development and benchmarking of automatic speech recognition (ASR) technologies in seven South African languages. Covering
This dataset contains tonal (F0/pitch) metadata extracted from dsfsi-anv/za-african-next-voices. zu
The AfVoices dataset is the largest open corpus of spontaneous Bambara speech at its release in late 2025. It contains 423 hours of segmented audio and 612 hours of original raw recordings collected across southern Mali. Speech was recorded in natural, conversation
The dataset was created by Digital Umuganda and made possible through funding from the Gates Foundation. The data spans five high-impact domains — Health, Government, Financial Services, Education, and Agriculture — to support robust ASR model development in both c