Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Vuk'uzenzele isiXhosa Speech Dataset (ViXSD)

Domain:

natural language processing

Record type:

dataset
Creator:
Lelapa AIVukosi Marivate

Vuk'uzenzele isiXhosa Speech Dataset (ViXSD) contains scripted narration of the Vuk’uzenzele South African Multilingual Corpus. ViXSD contains read speech from native speakers accompanied with rich metadata on speaker demographic and linguistic distribution.

ViXSD consists of 395 stereo audio recordings and corresponding transcriptions derived from the Vuk’uzenzele South African Multilingual Corpus.

It contains a total of 10 hours of narrated speech isiXhosa narrated by 8 speakers (4 male, 4 female) with approximately 39,000 words. We split the data into train, dev and test split for ease of use.

Visit

huggingface.copress release

Connected records

paperdataset

Tasks

speech processingautomatic speech recognition

Languages

Xhosa

Tags

vixsdlelapa aiway with words

Licenses

esethu license

Similar

Vuk'uzenzele isiXhosa Speech Dataset (ViXSD)

Vuk'uzenzele isiXhosa Speech Dataset (ViXSD)

Vuk'uzenzele isiXhosa Speech Dataset (ViXSD) contains scripted narration of the Vuk’uzenzele South A