This dataset comprises 2,488 high-quality audio recordings of read speech produced by a single Dagbani speaker over 16 sessions. Dagbani (ISO 639-3: dag), also known as Dagbane or Dagomba, is a Gur language of the Niger-Congo family spoken primarily in the Northern Region of Ghana, particularly in the Dagbon traditional area. It is the most widely spoken language in northern Ghana and serves as a lingua franca across the region. Despite being spoken by an estimated 3 to 4 million people, Dagbani remains severely under-resourced in terms of digital speech data, making this dataset a significant contribution to natural language processing efforts for the language.
Audio files are provided in MP3 format (approx. 185 MB), totalling 2 hours, 50 minutes and 59 seconds of speech. The dataset includes 16 audio/sentence mapping files in TSV format, containing 2,488 aligned audio/sentence pairs in total. Transcriptions follow the standard Dagbani orthography as developed by the Dagbani Orthography Committee and used in published Dagbani materials.
The recordings draw on a range of textual material in Dagbani, offering varied prosodic and lexical diversity for training and evaluating TTS and ASR models. The dataset is intended for research and scientific use in speech technology for Dagbani.