Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Tupuri-Bango_TTS-Dataset (female voice)

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Ins
Hôte:
Tupuri-Bango_TTS-Dataset (female voice) is a scripted speech dataset dedicated to the documentation and technological development of Tupuri (ISO 639-3: tui), a Chadic language of the Afro-Asiatic phylum spoken in the Kaele and Mayo-Danay Divisions of the Far-North Region of Cameroon and in adjacent areas of southern Chad. The dataset was compiled in the framework of the Mozilla Data Collective initiative (2026) and represents the Bango dialectal variety of Tupuri. Its principal added value lies in speaker gender: all 2,034 recordings are produced by a single female native speaker, in direct contrast to the previously released Sample Tupuri-Bango_TTS-Dataset (male voice) (Sample Tupuri-Bango_TTS-Dat…), which comprises exclusively male-voice recordings. Prior to this release, no female-voice speech resource existed for Tupuri-Bango. This dataset closes that gap, enabling the development of gender-balanced and multi-speaker TTS systems for the language, and making possible, for the first time, a direct comparison of male and female voice characteristics (pitch, prosody, tonal realisation) within the same dialectal variety and orthographic framework. The dataset comprises 2,034 high-quality MP3 audio recordings of Tupuri-Bango sentences read by the speaker across 21 recording sessions, together with per-session sentence-to-audio mapping files enabling precise alignment between textual and acoustic data. Sentences were drawn from a scripted speech prompt list and read in a controlled environment. The transcription of all sentences follows the General Alphabet of Cameroon's Languages (AGLC; French acronym: Alphabet Général des Langues Camerounaises), the reference standard for Cameroonian national languages. The Tupuri-Bango orthography employed in this dataset is characterised by an extended vowel inventory — including the open-mid front unrounded vowel ɛ and the open-mid back rounded vowel ɔ — a set of nasalized vowels represented by the tilde diacritic (ã, ẽ, ũ, õ), with long vowels encoded by vowel doubling (aa, ɛɛ, ɔɔ, ãã, etc.), a series of implosive consonants written with hooked letters (ɓ for the bilabial implosive and ɗ for the alveolar implosive), the velar nasal consonant ŋ (eng), a multi-register tone-marking system combining level (acute, grave) and contour (caron) diacritics applied to vowels, and the apostrophe (' or ') as a glottal stop or glottalization marker. The parallel availability of AGLC-transcribed text and aligned speech makes the dataset suitable for a wide range of applications, including text-to-speech (TTS) synthesis, automatic speech recognition (ASR), forced alignment, pronunciation modelling, and language learning tools. It also directly supports efforts to standardise and normalise the digital representation of Tupuri in language technology contexts.

Visit

mozilladatacollective.com

Tasks

text to speechspeech processing

Languages

BabangoMundangTupuri

Tags

mdcmozilla data collectiveTTSMP3TSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similaires

Sample Tupuri-Bango_TTS-DatasetBamun-TTS-Dataset (female voice)Tupuri-ASR-DatasetCommon Voice Scripted Speech 26.0 - Tupuri

Sample Tupuri-Bango_TTS-Dataset

Tupuri-bango_TTS-Dataset is a scripted speech dataset dedicated to the documentation and technologic

Bamun-TTS-Dataset (female voice)

This dataset comprises audio recordings of Bamun (Shupamem) speech aligned with textual transcriptio

Tupuri-ASR-Dataset

Tupuri-ASR-Dataset is a scripted speech dataset dedicated to the documentation and technological dev

Common Voice Scripted Speech 26.0 - Tupuri

A collection of read speech recordings in Tupuri (t'pur).