Logo Lanfrica

AwezaMed automatic speech recognition (ASR) test data

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Bandehorst, Jaco
Éditeur:
Van Niekerk, NinaCalteaux, Karen
Éditeur:
Voice Computing (VC) Research Group at the CSIR Nextgen Enterprises and Institutions (NGEI)
Hôte:avatar
The corpus contains orthographically transcribed broadband speech in four official languages of South Africa: Afrikaans, English, isiXhosa and isiZulu. Respondents read a number of ASR prompts (10 or 20) in a real-world environment. Dataset includes 1 hour of test data