Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

The South African Next Voices Multilingual Speech Dataset [Compressed]

Domain:

natural language processing

Record type:

dataset
Creator:
Marivate, VukosiOlaMunBak
Publisher:
Zenodo
Host:avatar

Swivuriso: ZA-African Next Voices-Compressed

Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 South African languages. The dataset is developed to support Automatic Speech Recognition (ASR) and inclusive speech technologies for low-resource African languages. It combines both scripted and unscripted speech, collected through ethical, community-centered processes.

Dataset Paper: ArXiv - Work in Progress

This is a compressed version of the original dataset Swivuriso: ZA-African Next…

⚠️ IMPORTANT: Visit the original dataset for full details

Language Coverage

LanguageTarget HoursReleased
isiZulu500▇▇▇▇▇▇▇▇▇▇ 100%
isiXhosa500▇▇▇▇▇▇▇▇▇▇ 100%
Sesotho500▇▇▇▇▇▇▇▇▇▇ 100%
Setswana500▇▇▇▇▇▇▇▇▇▇ 100%
Xitsonga500▇▇▇▇▇▇▇▇▇▇ 100%
isiNdebele250▇▇▇▇▇▇▇▇▇▇ 100%
Tshivenda250▇▇▇▇▇▇▇▇▇▇ 100%

Use Restriction:

The persons whose voices are included in this dataset, and the creators and owners of this dataset* do not give consent in any manner or form to, and strictly prohibit any use of this dataset for any form of text-to-speech (TTS), voice cloning, voice synthesis, or any technology or activity intended to replicate, mimic or generate human voices or any technology or activity resulting in the replication, mimicry or generation of human voices.

This dataset includes scripted and unscripted speech across various domains such as agriculture, health, finance, sports, transport, culture, society, and general topics. It is primarily designed for use in automatic speech recognition (ASR) tasks.

Use of this dataset for any form of text-to-speech (TTS), voice cloning, voice synthesis, or any technology intended to replicate or generate human voices is strictly prohibited.

These restrictions are in place until further notice.

Citations

If you use Swivuriso in your work, please cite both of the below:

 

Dataset

 
@dataset{za-african-next-voices-2025,
  title     = {The South African Next Voices Multilingual Speech Dataset},
    author       = {Marivate, Vukosi and
                  Olaleye, Kayode and
                  Mundia, Sitwala and
                  Bakainga, Andinda and
                  Netshifhefhe, Unarine Leo and
                  Milanzie, Mahmooda and
                  Mogale, Hope and
                  SINDANE, THAPELO and
                  Abdulrasaq, Zainab and
                  Mokgosi, Kesego and
                  Okorie, Chijioke and
                  van Wyk, Nia Zion and
                  Morrissey, Graham and
                  Dunbar, Dale and
                  Smit, Francois and
                  Chidi, Tsosheletso and
                  Mabuya, Rooweither and
                  Bukula, Andiswa and
                  MLAMBO, RESPECT and
                  Macucwa, Solomon Tebogo and
                  Abdulmumin, Idris and
                  Rananga, Seani},
  url2      = {github.com,
  url3      = {dsfsi.co.za,
  year      = {2025},
  type      = {dataset},
  publisher = {Zenodo},
  version   = {1.0},
  doi       = {10.5281/zenodo.17776289},
  url       = {doi.org,
}
 

https://huggingface.co/data…Research Paper

 

Will be available soon.


 

Visit

doi.org

Tasks

automatic speech recognitionspeech processing

Languages

NdebeleNdebeleSetswanaSotho, SouthernTsongaVendaXhosaZulu

Licenses

info:eu-repo/semantics/restrictedAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Swivuriso: The South African Next Voices Multilingual Speech DatasetAfrican Voices: Multilingual Speech Dataset for Low-Resource African Languagesza-african-next-voicesza-african-next-voicesMali African Next VoicesRwanda African Next Voices

Swivuriso: The South African Next Voices Multilingual Speech Dataset

This paper introduces Swivuriso, a 3000-hour multilingual speech dataset developed as part of the Af

African Voices: Multilingual Speech Dataset for Low-Resource African Languages

A large-scale multilingual speech dataset developed by Data Science Nigeria. Contains more than 3,000 hours of transcribed audio across four Nigerian languages: Hausa, Igbo, Nigerian Pidgin, and Yorùbá. The dataset supports Automatic Speech Recognition (ASR) and sp

za-african-next-voices

Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 S

za-african-next-voices

Note: This dataset is a compressed version of za-african-next-voices. It was compressed to .opus for

Mali African Next Voices

The AfVoices dataset is the largest open corpus of spontaneous Bambara speech at its release in late 2025. It contains 423 hours of segmented audio and 612 hours of original raw recordings collected across southern Mali. Speech was recorded in natural, conversation

Rwanda African Next Voices

The dataset was created by Digital Umuganda and made possible through funding from the Gates Foundation. The data spans five high-impact domains — Health, Government, Financial Services, Education, and Agriculture — to support robust ASR model development in both c