Logo Lanfrica

Sample-Ngomba-TTS-Dataset

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Ins
Hôte:
Sample-Ngomba-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technological development of Ngomba (ISO 639-3: jgo), a Grassfields Bantu language spoken in the Bamboutos Division of the West Region of Cameroon. The dataset was compiled in the framework of the Mozilla Data Collective initiative (2026). The dataset comprises 1033 high-quality audio recordings of Ngomba sentences read by a native speaker across 11 recording sessions (predominantly MP3 format, with 2 recordings in WAV format in session 07), together with per-session sentence-to-audio mapping files enabling precise alignment between textual and acoustic data. Sentences were drawn from a scripted speech prompt list and read in a controlled environment. The transcription of all sentences follows the General Alphabet of Cameroon's Languages (AGLC; French acronym: Alphabet Général des Langues du Cameroun), the reference standard for Cameroonian national languages. The Ngomba orthography employed in this dataset is distinguished by an extended vowel inventory — including the open-mid front unrounded vowel ɛ, the open-mid back rounded vowel ɔ, the high central rounded vowel ʉ (barred u), and the vowel ʉ̈ (barred u with diaeresis), which functions as a distinct phonemic grapheme in Ngomba — as well as a series of labialized consonants written by appending ẅ (w with diaeresis) to the base consonant (e.g., gẅ, sẅ, cẅ, kẅ, tsẅ), a multi-register tone-marking system combining level (acute, grave) and contour (caron, circumflex) diacritics applied to vowels and syllabic nasals, and the Latin small letter saltillo (ꞌ, U+A78C) for glottal closure. The parallel availability of AGLC-transcribed text and aligned speech makes the dataset suitable for a wide range of applications, including text-to-speech (TTS) synthesis, automatic speech recognition (ASR), forced alignment, pronunciation modelling, and language learning tools. It also directly supports efforts to standardise and normalise the digital representation of Ngomba in language technology contexts.