

Vuk'uzenzele isiXhosa Speech Dataset (ViXSD) contains scripted narration of the Vuk’uzenzele South African Multilingual Corpus. ViXSD contains read speech from native speakers accompanied with rich metadata on speaker demographic and linguistic distribution.
ViXSD consists of 395 stereo audio recordings and corresponding transcriptions derived from the Vuk’uzenzele South African Multilingual Corpus.
It contains a total of 10 hours of narrated speech isiXhosa narrated by 8 speakers (4 male, 4 female) with approximately 39,000 words. We split the data into train, dev and test split for ease of use.