Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Sample Medumba-TTS-Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
Ins
Host:
Medumba-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technological development of Medumba (ISO 639-3: byv), a Grassfields Bantu language spoken in the Ndé Division of the West Region of Cameroon. The dataset was compiled in the framework of the Mozilla Data Collective initiative (2026), in addition to the existing Common Voice Scripted Speech 25.0 – Medumba dataset (mozilladatacollective.com). The dataset comprises 994 high-quality MP3 audio recordings of Medumba sentences read by a native speaker across 10 recording sessions, together with per-session sentence-to-audio mapping files enabling precise alignment between textual and acoustic data. Sentences were drawn from a scripted speech prompt list and read in a controlled environment. The transcription of all sentences follows the General Alphabet of Cameroon's Languages (AGLC; French acronym: Alphabet Général des Langues Camerounaises), the reference standard for Cameroonian national languages. The Medumba AGLC orthography is distinguished by an extended vowel inventory — including the low back unrounded vowel ɑ, the open-mid front unrounded vowel ɛ, the open-mid back rounded vowel ɔ, the high central rounded vowel ʉ, and the mid central schwa ə — as well as a set of labialized consonants written by appending w to the base consonant (e.g., kw, gw, sw, bw), a series of pre-nasalized consonants written as digraphs or trigraphs (e.g., mb, nd, ŋg, ns, nsw), a two-level tone-marking system using grave (low) and contour diacritics (caron for rising LH; circumflex for falling HL) applied to vowels — high tone being unmarked — and the modifier letter apostrophe (ʼ) for the glottal stop. The parallel availability of AGLC-transcribed text and aligned speech makes the dataset suitable for a wide range of applications, including text-to-speech (TTS) synthesis, automatic speech recognition (ASR), forced alignment, pronunciation modelling, and language learning tools. It also directly supports efforts to standardise and normalise the digital representation of Medumba in language technology contexts.

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionspeech processingtext to speech

Languages

Medumba

Tags

mdcmozilla data collectiveTTSMP3TSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)