Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Bamun-TTS-Dataset (female voice)

Domain:

natural language processing

Record type:

dataset
Creator:
Ins
Host:
This dataset comprises audio recordings of Bamun (Shupamem) speech aligned with textual transcriptions. It is a female-voice companion to the previously published Bamun-TTS-Dataset (male voice), and is structured into 38 folders, each containing audio files and a corresponding audio-text mapping file. In total, the dataset contains 3,718 audio clips amounting to approximately 5 hours 4 minutes of speech. The audio clips are short, typically ranging from 1 to 10 seconds (with a small number of clips up to about 30 seconds), and are suitable for training and evaluating Text-to-Speech (TTS) systems. The dataset follows a structured format where each audio file is paired with its corresponding transcription in a tab-separated mapping file. The textual content used in this dataset originates from transcriptions of oral narratives documenting personal histories related to German colonisation in Cameroon. These texts were segmented into short utterances suitable for read speech and TTS modelling. The same textual material was used for the companion male-voice dataset.

Visit

mozilladatacollective.com

Tasks

speech processingtext to speech

Languages

Bamun

Tags

mdcmozilla data collectiveTTSWAVTSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similar

Bamun-TTS-Datasetmussacharles60/swahili-tts-female-voiceCommon Voice Scripted Speech 26.0 - BamunTupuri-Bango_TTS-Dataset (female voice)thisniyi/yoruba-tts-dataset-common-voice-single-36917TWB Voice Hausa TTS Dataset 1.0 - Sample Set

Bamun-TTS-Dataset

This dataset comprises audio recordings of Bamun (Shupamem) speech aligned with textual transcriptio

mussacharles60/swahili-tts-female-voice

Common Voice Scripted Speech 26.0 - Bamun

A collection of read speech recordings in Bamun (Shüpamom).

Tupuri-Bango_TTS-Dataset (female voice)

Tupuri-Bango_TTS-Dataset (female voice) is a scripted speech dataset dedicated to the documentation

thisniyi/yoruba-tts-dataset-common-voice-single-36917

TWB Voice Hausa TTS Dataset 1.0 - Sample Set

TWB Voice Hausa TTS 1.0 Sample Set is a high-quality text-to-speech corpus containing read speech data in Hausa, recorded by a single female speaker under acoustically optimal conditions. This dataset represents 10% of the complete Hausa TTS dataset collected as