This dataset comprises 1,521 high-quality audio recordings of read speech produced by a single Duala speaker over several sessions. Duala (ISO 639-3: dua), also known as Douala, is a Bantu language of the Niger-Congo family spoken primarily in the Littoral Region of Cameroon, notably in the city of Douala and its surrounding areas. It is a low-resource language with limited existing digital speech resources, making this dataset a significant contribution to natural language processing efforts for the language.
Audio files are provided in MP3 format (approx. 147 MB), totalling 4 hours, 30 minutes and 41.64 seconds of speech. The dataset includes 16 audio/sentence mapping files in TSV format, containing 1,521 aligned audio/sentence pairs in total. Transcriptions follow the General Alphabet of Cameroonian Languages, a standardised orthographic system based on the Latin alphabet augmented with phonetic characters and diacritical marks used to represent tonal and phonological features of Cameroonian languages.
The recordings draw on narrative texts relating to colonial encounters and experiences. These narratives originally existed as oral and audio recordings and were subsequently transcribed. The read-speech recordings therefore reflect a rich oral tradition rendered in text, offering valuable prosodic and lexical diversity for training and evaluating TTS and ASR models. The dataset is intended for research and scientific use in speech technology for Duala.