Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Basaa-ASR-Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
Ins
Host:
Basaa-ASR-Dataset is a curated speech dataset dedicated to the documentation and technological development of the Basaa language (ISO 639-3: bas), spoken in the Centre, Littoral, and South regions of Cameroon. The dataset was collected as part of a speech data initiative carried out at the École Normale Supérieure de Yaoundé (ENS-Yaoundé) in collaboration with the Mozilla Data Collective (MDC), and is intended to complement the two existing MDC datasets for Basaa: "Common Voice Scripted Speech 25.0 – Basaa" (mozilladatacollective.com) and "Common Voice Spontaneous Speech 3.0 – Basaa" (mozilladatacollective.com). All three datasets may be used complementarily depending on the task at hand. The added value of Basaa-ASR-Dataset is twofold. First, the transcription orthography used throughout the dataset is the General Alphabet of Cameroon's Languages (French acronym: AGLC), a standardised, phonologically motivated writing system that was insufficiently represented in previous Basaa speech datasets. Second, the dataset explicitly captures a broader range of Basaa dialectal varieties, including Basaa-ba-Yabassi and Babimbi, which were not sufficiently represented in earlier collections. The dataset comprises 2,498 MP3 audio recordings distributed across 25 recording sessions. Fifteen sessions (identified as asr-tts_dataset_bas_21 through bas_38) were conducted with 15 distinct speakers, each reading a set of 100 sentences in Basaa. The remaining 10 sessions (bassa_tts_dataset_01 through 10) constitute a dedicated TTS sub-corpus. All recordings were made using the MDC recording platform, which captures speaker metadata including the number of attempts per sentence. A total of 2,500 unique sentences were used across the dataset — 1,500 in the ASR/TTS sessions and 1,000 in the TTS sessions — with no overlap between the two sets. The total recorded audio amounts to 6,887 seconds (01:54:46).

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionspeech processing

Languages

BasaaBassa

Tags

mdcmozilla data collectiveASRTSVMP3

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)