Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Yoruba-TTS-Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
Ins
Host:
This dataset comprises audio recordings of Yoruba speech aligned with textual transcriptions. The dataset is structured into 17 folders, each containing audio files and a corresponding audio-text mapping file. The audio clips are short, typically ranging from 3 to 27 seconds, and are suitable for training and evaluating Text-to-Speech (TTS) systems. The dataset follows a structured format where each audio file is paired with its corresponding transcription in a tab-separated mapping file. The textual content used in this dataset originates from a variety of written and spoken sources in Yoruba, including narrative texts, conversational exchanges, opinion and commentary content, and everyday speech samples. These texts were segmented into short utterances suitable for read speech and TTS modelling.

Visit

mozilladatacollective.com

Tasks

speech processingtext to speech

Languages

Yoruba

Tags

mdcmozilla data collectiveTTSMP3TSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similar

Yoruba TTS DatasetPlotweaverModel/yoruba-tts-datasetthisniyi/yoruba-tts-dataset-openslrthisniyi/yoruba-tts-dataset-cv_and_openslrYoruba TTS (text-to-speech) training datasetthisniyi/yoruba-tts-dataset-openslr-single-7508

Yoruba TTS Dataset

A text-to-speech dataset for the Yoruba language. from datasets import load_dataset dataset = load

PlotweaverModel/yoruba-tts-dataset

thisniyi/yoruba-tts-dataset-openslr

thisniyi/yoruba-tts-dataset-cv_and_openslr

Yoruba TTS (text-to-speech) training dataset

Textbook audio archive size: total 36M archive created: 8 July 2011 mp3 file size ======== ==== 01-

thisniyi/yoruba-tts-dataset-openslr-single-7508