Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Sample Ngiemboon-TTS-Dataset

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Ins
Hôte:
Ngiemboon-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technological development of Ngiemboon (ISO 639-3: nnh), a Grassfields Bantu language spoken in the Bamboutos Division of the West Region of Cameroon. The dataset was compiled in the framework of the Mozilla Data Collective initiative (2026), as a supplement to the Common Voice Scripted Speech 25.0 – Ngiemboon dataset (mozilladatacollective.com). The dataset comprises 995 high-quality MP3 audio recordings of Ngiemboon sentences read by a native speaker across 10 recording sessions, together with per-session sentence-to-audio mapping files enabling precise alignment between textual and acoustic data. Sentences were drawn from a scripted speech prompt list and read in a controlled environment. The transcription of all sentences follows the General Alphabet of Cameroon's Languages (AGLC; French acronym: Alphabet Général des Langues Camerounaises), the reference standard for Cameroonian national languages. The Ngiemboon orthography employed in this dataset is distinguished by an extended vowel inventory — including the open-mid front unrounded vowel ɛ, the open-mid back rounded vowel ɔ, the high central rounded vowel ʉ, the high central unrounded vowel ɨ (barred i), and the close front rounded vowel ÿ — as well as a series of labialized consonants written by appending ẅ (w with diaeresis) to the base consonant (e.g., kẅ, gẅ, sẅ, zẅ, tsẅ), a multi-register tone-marking system combining level (acute, grave) and contour (caron, circumflex) diacritics applied to vowels and syllabic nasals, and the modifier letter apostrophe (ʼ) for glottal closure. The parallel availability of AGLC-transcribed text and aligned speech makes the dataset suitable for a wide range of applications, including text-to-speech (TTS) synthesis, automatic speech recognition (ASR), forced alignment, pronunciation modelling, and language learning tools. It also directly supports efforts to standardise and normalise the digital representation of Ngiemboon in language technology contexts.

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionspeech processingtext to speech

Languages

Ngiemboon

Tags

mdcmozilla data collectiveTTSMP3TSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similaires

Sample Dagbani-TTS-DatasetSample Mbo-TTS-DatasetSample Medumba-TTS-DatasetSample Fe’fe’-TTS-DatasetSample Batanga-TTS-DatasetSample-Ngomba-TTS-Dataset

Sample Dagbani-TTS-Dataset

This dataset comprises 2,488 high-quality audio recordings of read speech produced by a single Dagba

Sample Mbo-TTS-Dataset

Mbo-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technological develo

Sample Medumba-TTS-Dataset

Medumba-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technological de

Sample Fe’fe’-TTS-Dataset

Fe'fe'-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technological dev

Sample Batanga-TTS-Dataset

Batanga-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technological de

Sample-Ngomba-TTS-Dataset

Sample-Ngomba-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technologi