Previous studies have led to the collection of a considerable number of hours of open-source ASR data, for example, the work done in India where over 1000’s of hours of data were collected for low-resource Indian languages. In this research, we would like to transfer the learnings from these successes and replicate the same model for low-resource African languages. For example, the aspects around the use of speech data from noisy and non-noisy environments. However, to ensure that we proceed in a cost-efficient and sustainable approach, we deem it necessary to understand the amount of data that we need to collect for African languages. Hence, we propose to leverage the Mozilla Common Voice (MCV) platform and other appropriate and openly available / open-source repositories of African language datasets to build automatic speech recognition models and test their performance to learn if the data collected was sufficient. Platform overview A preview of what the platform contains and how to navigate. Use the links and tabs in the top navigation to jump to demos, datasets, results, or evaluation details. 1. Benchmark Datasets: A multilingual collection covering over 17 African languages, built from open corpora (e.g., Common Voice, Fleurs, NCHLT, ALFFA, Naija Voices). Each dataset is cleaned, validated, and partitioned into training, development, and test splits to ensure fair benchmarking.
Model Collections: Fine-tuned ASR models derived from Wav2Vec2 XLS-R, Whisper, MMS, and W2V-BERT, adapted for African phonetic, tonal, and orthographic features. These are hosted as public collections on Hugging Face.
Evaluation Scenarios: Designed to test data efficiency, domain adaptation, and speech-type robustness — e.g., how models generalize from read speech to spontaneous dialogue, or from education to agricultural domains.
ASR Demo Interface: A Gradio-powered live testing tool, allowing users to upload or record audio, view transcriptions, and submit structured feedback via the integrated backend API.
Quantitative Results: Comprehensive analysis of model performance across training hours and data scales (1–400 hours), visualized through Word Error Rate (WER) and Character Error Rate (CER) trends. Findings show clear data scaling laws, with XLS-R and W2V-BERT models performing best under low-resource conditions.
Human Evaluation Framework: A structured qualitative evaluation conducted with 20 native-language evaluators across 12 languages. Evaluators assessed accuracy, meaning preservation, orthography, and error types (e.g., named entities, punctuation, diacritics). This data is publicly available in the curated ASR_Evaluation_dataset.