Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Sample Dagbani-TTS-Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
Ins
Host:
This dataset comprises 2,488 high-quality audio recordings of read speech produced by a single Dagbani speaker over 16 sessions. Dagbani (ISO 639-3: dag), also known as Dagbane or Dagomba, is a Gur language of the Niger-Congo family spoken primarily in the Northern Region of Ghana, particularly in the Dagbon traditional area. It is the most widely spoken language in northern Ghana and serves as a lingua franca across the region. Despite being spoken by an estimated 3 to 4 million people, Dagbani remains severely under-resourced in terms of digital speech data, making this dataset a significant contribution to natural language processing efforts for the language. Audio files are provided in MP3 format (approx. 185 MB), totalling 2 hours, 50 minutes and 59 seconds of speech. The dataset includes 16 audio/sentence mapping files in TSV format, containing 2,488 aligned audio/sentence pairs in total. Transcriptions follow the standard Dagbani orthography as developed by the Dagbani Orthography Committee and used in published Dagbani materials. The recordings draw on a range of textual material in Dagbani, offering varied prosodic and lexical diversity for training and evaluating TTS and ASR models. The dataset is intended for research and scientific use in speech technology for Dagbani.

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionspeech processingtext to speech

Languages

Dagbani

Tags

mdcmozilla data collectiveTTSMP3TSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similar

Sample Fe’fe’-TTS-DatasetSample Batanga-TTS-DatasetSample Mbo-TTS-DatasetSample-Ngomba-TTS-DatasetSample Medumba-TTS-DatasetSample Ngiemboon-TTS-Dataset

Sample Fe’fe’-TTS-Dataset

Fe'fe'-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technological dev

Sample Batanga-TTS-Dataset

Batanga-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technological de

Sample Mbo-TTS-Dataset

Mbo-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technological develo

Sample-Ngomba-TTS-Dataset

Sample-Ngomba-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technologi

Sample Medumba-TTS-Dataset

Medumba-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technological de

Sample Ngiemboon-TTS-Dataset

Ngiemboon-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technological