Eton-ASR-Dataset is a scripted speech dataset dedicated to the documentation and technological development of Eton (ISO 639-3: etn), a Narrow Bantu language spoken primarily in the Centre Region of Cameroon. The dataset was compiled at the École Normale Supérieure of Yaoundé, in the department of Cameroonian languages and Cultures (2026).
The dataset comprises 1,402 high-quality MP3 audio recordings of Eton sentences read by 11 native speakers across 15 recording sessions, together with per-session sentence-to-audio mapping files enabling precise alignment between textual and acoustic data. Sentences were drawn from a scripted speech prompt list and read by each speaker in a controlled environment.
The primary added value of this dataset lies in its orthographic alignment with the General Alphabet of Cameroon's Languages (AGLC; French acronym: AGLC — Alphabet Général des Langues Camerounaises), the reference standard for Cameroonian national languages. In particular, this dataset preserves systematic tone marking, a feature that the existing Common Voice Scripted Speech 25.0 – Eton dataset available on the Mozilla Data Collective platform tends to omit. By making tone information explicit in the transcription, this dataset enables the development and evaluation of speech technology models that are sensitive to the tonal contrasts that are phonemically contrastive in Eton.
From a methodological perspective, the dataset is designed to complement the existing Common Voice Scripted Speech resource for Eton rather than to replace it, thereby extending the total amount of available Eton speech data aligned with an orthographically principled transcription standard. The parallel availability of AGLC-transcribed text and aligned speech makes the dataset suitable for a wide range of applications, including automatic speech recognition (ASR), text-to-speech (TTS), forced alignment, pronunciation modelling and language learning tools. It also directly supports efforts to standardise and normalise the digital representation of Eton in language technology contexts.