Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Nigerian Common Voice Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
ben
Host:
The Nigerian Common Voice Dataset is a comprehensive dataset consisting of 158 hours of audio recordings and corresponding transcription (sentence). This dataset includes metadata like accent, locale that can help improve the accuracy of speech recognition engines. This dataset is specifically curated to address the gap in speech and language

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Languages

HausaIgboYoruba

Licenses

apache-2.0

Similar

Nigerian Common Voice DatasetmbazaNLP/common-voice-kinyarwanda-english-datasetCommon Voice Scripted Speech 26.0 - Nigerian Pidgin Englishthisniyi/yoruba-tts-dataset-common-voice-single-36917Common VoiceCommon Voice Basaa

Nigerian Common Voice Dataset

The Nigerian Common Voice Dataset is a comprehensive dataset consisting of 158 hours of audio record

mbazaNLP/common-voice-kinyarwanda-english-dataset

A compilation of Kinyarwanda-english dataset to be used to train multi-lingual ASR Note: The audio d

Common Voice Scripted Speech 26.0 - Nigerian Pidgin English

A collection of read speech recordings in Nigerian Pidgin English (pcm).

thisniyi/yoruba-tts-dataset-common-voice-single-36917

Common Voice

Common Voice is Mozilla's initiative to help teach machines how real people speak. The dataset currently consists of 7,335 validated hours of speech in 60 languages, but we’re always adding more voices and languages.

Common Voice Basaa

Voice data collection and distribution Interface for Basaa language using Mozilla's Common Voice infrastructure.