This dataset comprises audio recordings of Bamun (Shupamem) speech aligned with textual transcriptions. It is a female-voice companion to the previously published Bamun-TTS-Dataset (male voice), and is structured into 38 folders, each containing audio files and a corresponding audio-text mapping file. In total, the dataset contains 3,718 audio clips amounting to approximately 5 hours 4 minutes of speech.
The audio clips are short, typically ranging from 1 to 10 seconds (with a small number of clips up to about 30 seconds), and are suitable for training and evaluating Text-to-Speech (TTS) systems. The dataset follows a structured format where each audio file is paired with its corresponding transcription in a tab-separated mapping file.
The textual content used in this dataset originates from transcriptions of oral narratives documenting personal histories related to German colonisation in Cameroon. These texts were segmented into short utterances suitable for read speech and TTS modelling. The same textual material was used for the companion male-voice dataset.