Logo Lanfrica

African Voices Nigeria: 2500 hours of ethically sourced speech data for four Nigerian Languages

Domaine:

natural language processing

Type de record:

datasetpaper
Créateur:
AssAdeAdeAnj
Éditeur:
Und
Hôte:avatar
African languages remain severely underrepresented in large-scale speech resources, particularly for spontaneous, naturally occurring speech that reflects real-world linguistic use. We present African Voices, a large-scale, ethically governed speech dataset covering four Nigerian languages, [~2,500] hours of audio, and 2,865 speakers, with a focus on spontaneous and scripted speech across diverse sociolinguistic contexts. Unlike existing resources that primarily rely on read or scripted speech, African Voices captures natural variation in accent, dialect, register, and code-switching, accompanied by rich demographic and contextual metadata. We describe the data collection methodology, transcription and a principled governance framework designed to support responsible use of speech data in low-resource settings. We further provide baseline automatic speech recognition results across languages. African Voices enables research on robust and fair ASR and serves as a foundational resource for advancing NLP research in African languages.

Similaires