This dataset comprises audio recordings of isiXhosa speech aligned with textual transcriptions. The dataset is structured into 24 folders, each containing audio files and a corresponding audio-text mapping file.
The audio clips are short, typically ranging from 1 to 13 seconds, and are suitable for training and evaluating Text-to-Speech (TTS) systems. The dataset follows a structured format where each audio file is paired with its corresponding transcription in a tab-separated mapping file.
The textual content used in this dataset originates from written isiXhosa sources published on the indigenous-language blogging platform IndigenousBlogs (
indigenousblogs.com), which hosts original content authored by isiXhosa-speaking bloggers across a range of topics, including narrative texts, opinion pieces, cultural commentary, and everyday informational content. These texts were segmented into short utterances suitable for read speech and TTS modelling.