Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

WA-Vox: Building a Large Scale, Inclusive Speech Dataset for West African Languages

Domain:

natural language processing

Record type:

dataset
Creator:
Odu
Publisher:
Zenodo
Host:avatar
Empowering West African Voices: Introducing WA-VOX West Africa, home to hundreds of millions across vibrant linguistic diversity, remains sidelined by speech AI that prioritizes English and Mandarin. WA-VOX changes that by assembling a robust, 850-hour corpus of scripted and unscripted speech capturing everyday conversations, market chatter, and formal discussions in languages like Hausa, Yoruba, Zulu etc. This isn't just data, it's a call to action for equitable tech, built through open collaborations with local universities and communities in Nigeria, Ghana, Senegal, Mali, and Burkina Faso. Why It Matters Current datasets like Common Voice offer mere scraps (under 20 hours for Hausa), limiting real world ASR tools and perpetuating access barriers. WA-VOX fills this void with deep domain coverage (agriculture to culture), dialectal variety, and rigorous quality controls. As an independent researcher, I've designed it for sustainability: free for non commercial use, with a Creative Commons license and ongoing community stewardship. Next Steps and Access This document outlines the full roadmap from recruitment of 1,200 diverse speakers to risk mitigated collection and open release. Download to explore methodologies, baselines, and ethics protocols. Join the movement: collaborate, cite, or fund to power the next wave of inclusive language tech. Questions? Reach out via royalrubee@gmail.com.

Visit

doi.orgzenodo.org

Tasks

automatic speech recognitionspeech processing

Languages

HausaYoruba

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African LanguagesThe Thiomi Dataset: A Large-Scale Multimodal Corpus for Low-Resource African LanguagesWAXAL: A Large-Scale Multilingual African Language Speech CorpusPolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and DialectsOpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource LanguagesTowards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages

The development of high-performing, robust, and reliable speech technologies depends on large, high-quality datasets. However, African languages -- including our focus, Igbo, Hausa, and Yoruba -- remain under-represented due to insufficient data. Popular voice-e

The Thiomi Dataset: A Large-Scale Multimodal Corpus for Low-Resource African Languages

We present the Thiomi Dataset, a large-scale multimodal corpus spanning ten African languages across

WAXAL: A Large-Scale Multilingual African Language Speech Corpus

The advancement of speech technology has predominantly favored high-resource languages, creating a significant digital divide for speakers of most Sub-Saharan African languages. To address this gap, we introduce WAXAL, a large-scale, openly accessible speech datase

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

While End-to-End (E2E) Speech-Large Language Models (Speech-LLMs) are rapidly evolving, their evalua

OpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource Languages

Recent advances in neural text-to-speech (TTS) and multilingual speech generation have substantially

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages

Speech large language models (SLLMs) built on speech encoders, adapters, and LLMs demonstrate remark