Tupuri-bango_TTS-Dataset is a scripted speech dataset dedicated to the documentation and technological development of Tupuri (ISO 639-3: tui), a Chadic language of the Afro-Asiatic phylum spoken in the Kaele and Mayo-Danay Divisions of the Far-North Region of Cameroon and in adjacent areas of southern Chad. The dataset was compiled in the framework of the Mozilla Data Collective initiative (2026) as a supplement to the Common Voice Scripted Speech 25.0 – Tupuri dataset (mozilladatacollective.com) and to the Tupuri-ASR-Dataset (Tupuri-ASR-Dataset). The speech material represents the Bango dialectal variety of Tupuri.
The dataset comprises 983 high-quality MP3 audio recordings of Tupuri-Bango sentences read by a native speaker across 10 recording sessions, together with per-session sentence-to-audio mapping files enabling precise alignment between textual and acoustic data. Sentences were drawn from a scripted speech prompt list and read in a controlled environment.
The transcription of all sentences follows the General Alphabet of Cameroon's Languages (AGLC; French acronym: Alphabet Général des Langues Camerounaises), the reference standard for Cameroonian national languages. The Tupuri-Bango orthography employed in this dataset is characterised by an extended vowel inventory — including the open-mid front unrounded vowel ɛ and the open-mid back rounded vowel ɔ — a set of nasalized vowels represented by the tilde diacritic (ã, ẽ, ũ, õ), with long vowels encoded by vowel doubling (aa, ɛɛ, ɔɔ, ãã, etc.), a series of implosive consonants written with hooked letters (ɓ for the bilabial implosive and ɗ for the alveolar implosive), the velar nasal consonant ŋ (eng), a multi-register tone-marking system combining level (acute, grave) and contour (caron) diacritics applied to vowels, and the apostrophe (', ') as a glottal stop or glottalization marker. The parallel availability of AGLC-transcribed text and aligned speech makes the dataset suitable for a wide range of applications, including text-to-speech (TTS) synthesis, automatic speech recognition (ASR), forced alignment, pronunciation modelling, and language learning tools. It also directly supports efforts to standardise and normalise the digital representation of Tupuri in language technology contexts.