Logo Lanfrica

AwezaMed automatic speech recognition (ASR) test data

Domain:

natural language processing

Record type:

dataset
Creator:
Bandehorst, Jaco
Editor:
Van Niekerk, NinaCalteaux, Karen
Publisher:
Voice Computing (VC) Research Group at the CSIR Nextgen Enterprises and Institutions (NGEI)
Host:avatar
The corpus contains orthographically transcribed broadband speech in four official languages of South Africa: Afrikaans, English, isiXhosa and isiZulu. Respondents read a number of ASR prompts (10 or 20) in a real-world environment. Dataset includes 1 hour of test data