Batanga-TTS-Dataset is a scripted speech dataset dedicated to the documentation and technological development of Batanga (ISO 639-3: bnm), a Bantu language spoken along the Atlantic coast of the Ocean Division, South Region of Cameroon. The dataset was compiled in the framework of the Mozilla Data Collective initiative (2026), as a supplement to the Common Voice Scripted Speech 25.0 – Batanga dataset (mozilladatacollective.com) and the Tupuri-ALCAM-MultimodalDataset (Batanga-ALCAM-MultimodalDat…).
The dataset comprises 1,023 high-quality MP3 audio recordings of Batanga sentences read by a native speaker across 11 recording sessions, together with per-session sentence-to-audio mapping files enabling precise alignment between textual and acoustic data. Sentences were drawn from a scripted speech prompt list and read in a controlled environment. The dataset represents both principal speech varieties of Batanga — Bapuku and Banoho (also spelt Banoo or Banɔɔ).
The transcription of all sentences follows the General Alphabet of Cameroon's Languages (AGLC; French acronym: Alphabet Général des Langues Camerounaises), the reference standard for Cameroonian national languages. The Batanga orthography employed in this dataset is distinguished by an extended vowel inventory — including the open-mid front unrounded vowel ɛ and the open-mid back rounded vowel ɔ — as well as a tone-marking system using acute and grave diacritics applied to vowels, nasal-consonant boundary markers (nʼ, ŋʼ) separating nasal prefixes from consonant clusters, and the eng symbol (ŋ) for the velar nasal. The parallel availability of AGLC-transcribed text and aligned speech makes the dataset suitable for a wide range of applications, including text-to-speech (TTS) synthesis, automatic speech recognition (ASR), forced alignment, pronunciation modelling, and language learning tools. It also directly supports efforts to standardise and normalise the digital representation of Batanga in language technology contexts.