Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Bulu-ASR-Dataset

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Ins
Hôte:
Bulu-ASR-Dataset is a scripted speech dataset dedicated to the documentation and technological development of Bulu (ISO 639-3: bum), a Narrow Bantu language spoken primarily in the South and Centre Regions of Cameroon. The dataset was compiled at the École Normale Supérieure de Yaoundé (2026). The dataset comprises 819 high-quality MP3 audio recordings of Bulu sentences read by 8 native speakers across 9 recording sessions, together with per-session sentence-to-audio mapping files enabling precise alignment between textual and acoustic data. Sentences were drawn from a scripted speech prompt list and read by each speaker in a controlled environment. The primary added value of this dataset lies in its demographic composition: the majority of contributing speakers are female, a demographic group that is significantly underrepresented in the existing Common Voice Scripted Speech 25.0 – Bulu dataset available on the Mozilla Data Collective platform. By providing a substantial body of high-quality female Bulu speech, this dataset directly addresses the speaker gender imbalance in available Bulu speech resources and enables the development and evaluation of speech technology models that are more robust across speaker genders. The dataset follows the orthography established by the American Presbyterian Mission (Mission Protestante Américaine, MPA), the historically grounded and community-recognised writing standard for Bulu. This orthography, developed by MPA missionaries and Bulu-speaking collaborators from the late nineteenth century onwards, was codified through the Bulu Bible translation, grammar descriptions, and literacy materials that have shaped Bulu literacy for over a century. From a methodological perspective, the dataset is designed to complement the existing Common Voice Scripted Speech resource for Bulu rather than to replace it, thereby extending the total amount of available Bulu speech data while improving demographic coverage and orthographic fidelity. The parallel availability of MPA-transcribed text and aligned speech makes the dataset suitable for a wide range of applications, including automatic speech recognition (ASR), text-to-speech (TTS), forced alignment, pronunciation modelling and language learning tools.

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionspeech processingtext to speech

Languages

Bulu

Tags

mdcmozilla data collectiveASRMP3TSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)