Logo Lanfrica

IsiZulu Second Language Learner Speech Corpus

Domain:

natural language processingeducation

Record type:

dataset
Creator:
O'Neil, AlexandraHjortnaes, NilsNkosi, ZinhleNdlovu, Thulile
Publisher:
Indiana University
Host:avatar
This corpus is specifically designed to assist in evaluating the performance of pronunciation feedback tools for second language learning. The corpus is comprised of gold standard recordings from isiZulu teachers (2,493 recordings) and recordings from isiZulu L2 learners that have been annotated by isiZulu teachers for phonemic and tonal pronunciation errors (9,639 recordings). The accompanying database and tsv file include the teacher annotations and demographic information.