This dataset comprises 1,737 high-quality audio recordings of read speech produced by a single Afaan Oromo speaker over 24 sessions. Afaan Oromo (ISO 639-3: orm), also known as Oromo or Oromiffa, is a Cushitic language of the Afroasiatic family and the most widely spoken language in Ethiopia. It is spoken primarily in the Oromia region of Ethiopia, with significant speaker communities in neighbouring regions and in Kenya, Somalia, and the diaspora. Despite being one of the most widely spoken languages in Africa — with an estimated 40 to 50 million speakers — it remains severely under-resourced in terms of digital speech data, making this dataset a significant contribution to natural language processing efforts for the language.
Audio files are provided in MP3 format (approx. 266 MB), totalling 4 hours, 51 minutes and 7 seconds of speech. The dataset includes 24 audio/sentence mapping files in TSV format, containing 1,737 aligned audio/sentence pairs in total. Transcriptions follow the Qubee orthographic system, the standardised Latin-based alphabet officially adopted for Afaan Oromo in 1991.
The recordings draw on news reports and biographical narratives in Afaan Oromo. These texts reflect contemporary journalistic and narrative registers of the language, offering varied prosodic and lexical diversity for training and evaluating TTS and ASR models. The dataset is intended for research and scientific use in speech technology for Afaan Oromo.