Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

CSIR SAMA Speech Corpus Manual Datasets

Domain:

natural language processing

Record type:

dataset
Creator:
Bandehorst, JacoMak, Franco
Publisher:
Voice Computing (VC) Research Group at the CSIR Nextgen Enterprises and Institutions (NGEI)SADiLaR
Host:avatar
The evaluation corpus contains orthographically transcribed broadband speech in Afrikaans, isiXhosa, isiZulu, Sepedi, Sesotho, Tshivenḓa all part of South Africa’s eleven official written languages. The audio was harvested as MP3 podcasts and automatically segmented and transcribed. Segment transcriptions are provided in XML format.
Afrikaans
• News-20:04:15-Bulletins:377
• Drama-12:35:47-Episodes: 351
Sepedi
• Drama-10:25:16-Episodes: 321
Sesotho
• News-14:49:11-Bulletins:326
• Drama-10:01:17-Episodes: 200
isiXhosa
• News-14:57:20-Bulletins: 325
• Drama-09:41:42-Episodes: 190
isiZulu
• News-15:55:14-Bulletins:349
• Drama-07:58:10-Episodes: 124
Tshivenḓa
• Drama-12:32:29-Episodes: 275

Visit

hdl.handle.net

Tasks

automatic speech recognitionspeech processing

Languages

AfrikaansSotho, NorthernSotho, SouthernVendaXhosaZulu

Tags

Speech corporaData harvestingTranscriptionSegmentation

Licenses

Creative Commons Attribution 4.0 International (CC-BY 4.0): https://www.creativecommons.org/licenses/by/4.0/

Similar

Faouzielbakri/DARIJA-CORPUS-COMMENTS-DATASETSHausa Speech CorpusAmharic Speech CorpusTshivenda Speech CorpusSPCS Speech CorpusNCHLT speech corpus

Faouzielbakri/DARIJA-CORPUS-COMMENTS-DATASETS

Hausa Speech Corpus

This is a Hausa Speech data set that was recorded as a baseline for Hausa Speech Recognition. The data sets can be used in building Automatic Speech recognition for Hausa language, Speech synthesis and speaker recognition.

Amharic Speech Corpus

This is an Amharic speech corpus which is suitable for the development and evaluation of speech recognition and retrieval systems. The corpus contains 110 hours of speech data with syllable and grapheme-based transcriptions collected from public domain or resources

Tshivenda Speech Corpus

A telephone speech database for Tshivenda ASR system was created from transcriptions of speech data

SPCS Speech Corpus

Broadband speech corpus of approximately 10 hours and the corresponding transcriptions. The devel

NCHLT speech corpus

The NCHLT speech corpus of the South African languages