This dataset comprises audio recordings of Tiv speech aligned with textual transcriptions. The dataset is structured into 14 folders, each containing audio files and a corresponding audio-text mapping file.
The audio clips are short, typically ranging from 1 to 30 seconds, and are suitable for training and evaluating Text-to-Speech (TTS) systems. The dataset follows a structured format where each audio file is paired with its corresponding transcription in a tab-separated mapping file.
The textual content used in this dataset originates from Tiv namel, selected traditional oral-narrative material that has been committed to writing (folk tales, proverbs, riddles, and cultural/ethnographic accounts). These texts were segmented into short utterances suitable for read speech and TTS modelling.