Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Yoruba Speech Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
Sil
Host:
The most comprehensive Yoruba speech dataset on HuggingFace - natural, real-world Yoruba from native speakers in Nigeria and the diaspora. Total audio samples: 39 recordings Total duration: ~22 minutes Primary region: Nigeria (Southwest - Ibadan, Lagos) Context: Natural spontaneous speech (free_speech) Audio format: WAV files Sample rate: 48 kHz License: CC BY-NC 4.0 (free for research, non-commercial use)

Visit

huggingface.co

Languages

Yoruba

Tags

yorubanigerian-languageswest-africanigeriatonal-languageafrican-languageslow-resourcespeech-datavoice-aiasr+1

Licenses

cc-by-nc-4.0

Similar

yoruba speech datasetshunyalabs/yoruba-speech-datasetYoruba Speech-Text Parallel DatasetYoruba Speech-Text Parallel DatasetYoruba TTS (text-to-speech) training datasetcarlshizi/yoruba-speech-pipelines

yoruba speech dataset

Yoruba Language Audio Dataset (12 Speakers, 6 Hours of Recordings)

shunyalabs/yoruba-speech-dataset

Yoruba Speech-Text Parallel Dataset

Speech recognition, text-to-speech synthesis, voice assistants, language modeling Notes / challenge

Yoruba Speech-Text Parallel Dataset

This dataset contains 1647022 parallel speech-text pairs for Yoruba, a language spoken primarily in

Yoruba TTS (text-to-speech) training dataset

Textbook audio archive size: total 36M archive created: 8 July 2011 mp3 file size ======== ==== 01-

carlshizi/yoruba-speech-pipelines

A research writing sample proposing a data-centric speech pipeline for low-resource Yoruba ASR, TTS,