Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Eton-TTS-Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
Ins
Host:
Eton-TTS-Dataset is a single-speaker scripted speech dataset dedicated to the documentation and technological development of Eton (ISO 639-3: eto), a Narrow Bantu language spoken primarily in the Centre Region of Cameroon. The dataset was compiled at the École Normale Supérieure de Yaoundé in the framework of the Mozilla Data Collective (2026). The dataset comprises 1,802 audio clips of Eton sentences read by a single native female speaker across 19 recording sessions (2026-06-18 to 2026-07-01), one per source text, together with per-session sentence-to-audio mapping files enabling precise alignment between textual and acoustic data. Sentences were drawn from the same scripted speech prompt list used for the companion Eton-ASR-Dataset and read in a controlled environment by one consistent voice, which is the design requirement for training or fine-tuning corpus-based TTS voice models. As with the companion Eton-ASR-Dataset, the primary added value of this dataset lies in its orthographic alignment with the General Alphabet of Cameroon's Languages (AGLC; French acronym: AGLC — Alphabet Général des Langues Camerounaises), the reference standard for Cameroonian national languages. In particular, the dataset preserves systematic tone marking, a feature that the existing Common Voice Scripted Speech 25.0 – Eton dataset available on the Mozilla Data Collective platform tends to omit. By making tone information explicit in the transcription, this dataset enables the development and evaluation of speech synthesis models that are sensitive to the tonal contrasts that are phonemically contrastive in Eton. From a methodological perspective, this dataset complements the multi-speaker Eton-ASR-Dataset: while the latter is intended for ASR robustness across speakers, Eton-TTS-Dataset provides a single consistent voice reading near-complete coverage of the same 19-text prompt list (1,802 of the available scripted sentences), making it directly suited to building a single-voice Eton TTS system aligned with AGLC orthography and tone marking.

Visit

mozilladatacollective.com

Tasks

speech processingtext to speech

Languages

Eton

Tags

mdcmozilla data collectiveTTSWAVTSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similar

Eton-ASR-DatasetAn Eton Sociocultural DatasetSuundi-TTS-DatasetRw Tts DatasetBomitaba-TTS-DatasetBeembe-TTS-Dataset

Eton-ASR-Dataset

Eton-ASR-Dataset is a scripted speech dataset dedicated to the documentation and technological devel

An Eton Sociocultural Dataset

A Eton Sociocultural Dataset is a sociocultural and onomastic dataset documenting the Eton community

Suundi-TTS-Dataset

The dataset consists of paired audio and text data on Suundi (sdj), a language spoken in Congo. The

Rw Tts Dataset

Kinyarwanda (rw) text-to-speech dataset. Studio-recorded read speech aligned with transcriptions, co

Bomitaba-TTS-Dataset

The dataset comprises three components: audio clips, an audio mapping file, and raw audio of Bomitab

Beembe-TTS-Dataset

The dataset consists of paired audio and text data on Beembe (beq), a language spoken in Congo. The