Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Swivuriso: ZA-African Next Voices

Domain:

agricultureeducationsocioeconomic

Record type:

datasetpaper
Creator:
DSFSI

Swivuriso is a 3000-hour multilingual speech dataset developed as part of the African Next Voices project, to support the development and benchmarking of automatic speech recognition (ASR) technologies in seven South African languages. Covering agriculture, healthcare, and general domain topics, Swivuriso addresses significant gaps in existing ASR datasets.

Visit

huggingface.coarxiv.orgwww.dsfsi.co.za

Tasks

automatic speech recognitionspeech processing

Languages

NdebeleSetswanaSotho, SouthernTsongaVendaXhosaZulu

Tags

African Next VoicesANVASRSouth AfricaSwivuriso

Licenses

CC BY 4.0

Similar

za-african-next-voicesza-african-next-voicesSwivuriso: The South African Next Voices Multilingual Speech Datasetkesbeast23/za-african-next-voices-tonalkesbeast23/za-african-next-voices-difficulty-scoresAfrican Next Voices: Ethiopia

za-african-next-voices

Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 S

za-african-next-voices

Note: This dataset is a compressed version of za-african-next-voices. It was compressed to .opus for

Swivuriso: The South African Next Voices Multilingual Speech Dataset

This paper introduces Swivuriso, a 3000-hour multilingual speech dataset developed as part of the Af

kesbeast23/za-african-next-voices-tonal

This dataset contains tonal (F0/pitch) metadata extracted from dsfsi-anv/za-african-next-voices. zu

kesbeast23/za-african-next-voices-difficulty-scores

African Next Voices: Ethiopia

AfriVoice Ethiopia is an open-source speech corpus for ASR development covering five Ethiopian languages: Amharic, Afaan Oromo, Sidama, Wolaytta, and Tigrinya.