Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Sample Medumba-TTS-Dataset

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Ins
Hôte:
Medumba-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technological development of Medumba (ISO 639-3: byv), a Grassfields Bantu language spoken in the Ndé Division of the West Region of Cameroon. The dataset was compiled in the framework of the Mozilla Data Collective initiative (2026), in addition to the existing Common Voice Scripted Speech 25.0 – Medumba dataset (mozilladatacollective.com). The dataset comprises 994 high-quality MP3 audio recordings of Medumba sentences read by a native speaker across 10 recording sessions, together with per-session sentence-to-audio mapping files enabling precise alignment between textual and acoustic data. Sentences were drawn from a scripted speech prompt list and read in a controlled environment. The transcription of all sentences follows the General Alphabet of Cameroon's Languages (AGLC; French acronym: Alphabet Général des Langues Camerounaises), the reference standard for Cameroonian national languages. The Medumba AGLC orthography is distinguished by an extended vowel inventory — including the low back unrounded vowel ɑ, the open-mid front unrounded vowel ɛ, the open-mid back rounded vowel ɔ, the high central rounded vowel ʉ, and the mid central schwa ə — as well as a set of labialized consonants written by appending w to the base consonant (e.g., kw, gw, sw, bw), a series of pre-nasalized consonants written as digraphs or trigraphs (e.g., mb, nd, ŋg, ns, nsw), a two-level tone-marking system using grave (low) and contour diacritics (caron for rising LH; circumflex for falling HL) applied to vowels — high tone being unmarked — and the modifier letter apostrophe (ʼ) for the glottal stop. The parallel availability of AGLC-transcribed text and aligned speech makes the dataset suitable for a wide range of applications, including text-to-speech (TTS) synthesis, automatic speech recognition (ASR), forced alignment, pronunciation modelling, and language learning tools. It also directly supports efforts to standardise and normalise the digital representation of Medumba in language technology contexts.

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionspeech processingtext to speech

Languages

Medumba

Tags

mdcmozilla data collectiveTTSMP3TSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similaires

Sample Ngiemboon-TTS-DatasetSample Fe’fe’-TTS-DatasetSample Mbo-TTS-DatasetSample Batanga-TTS-DatasetSample-Ngomba-TTS-DatasetSample Dagbani-TTS-Dataset

Sample Ngiemboon-TTS-Dataset

Ngiemboon-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technological

Sample Fe’fe’-TTS-Dataset

Fe'fe'-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technological dev

Sample Mbo-TTS-Dataset

Mbo-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technological develo

Sample Batanga-TTS-Dataset

Batanga-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technological de

Sample-Ngomba-TTS-Dataset

Sample-Ngomba-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technologi

Sample Dagbani-TTS-Dataset

This dataset comprises 2,488 high-quality audio recordings of read speech produced by a single Dagba