Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Swivuriso: ZA-African Next Voices

Domaine:

agricultureeducationsocioeconomic

Type de record:

datasetpaper
Créateur:
DSFSI

Swivuriso is a 3000-hour multilingual speech dataset developed as part of the African Next Voices project, to support the development and benchmarking of automatic speech recognition (ASR) technologies in seven South African languages. Covering agriculture, healthcare, and general domain topics, Swivuriso addresses significant gaps in existing ASR datasets.

Visit

huggingface.coarxiv.orgwww.dsfsi.co.za

Tasks

automatic speech recognitionspeech processing

Languages

NdebeleSetswanaSotho, SouthernTsongaVendaXhosaZulu

Tags

African Next VoicesANVASRSouth AfricaSwivuriso

Licenses

CC BY 4.0

Similaires

za-african-next-voicesza-african-next-voicesSwivuriso: The South African Next Voices Multilingual Speech Datasetkesbeast23/za-african-next-voices-tonalkesbeast23/za-african-next-voices-difficulty-scoresMali African Next Voices

za-african-next-voices

Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 S

za-african-next-voices

Note: This dataset is a compressed version of za-african-next-voices. It was compressed to .opus for

Swivuriso: The South African Next Voices Multilingual Speech Dataset

This paper introduces Swivuriso, a 3000-hour multilingual speech dataset developed as part of the Af

kesbeast23/za-african-next-voices-tonal

This dataset contains tonal (F0/pitch) metadata extracted from dsfsi-anv/za-african-next-voices. zu

kesbeast23/za-african-next-voices-difficulty-scores

Mali African Next Voices

The AfVoices dataset is the largest open corpus of spontaneous Bambara speech at its release in late 2025. It contains 423 hours of segmented audio and 612 hours of original raw recordings collected across southern Mali. Speech was recorded in natural, conversation