Logo Lanfrica

Yoruba Speech-Text Parallel Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
mic
Host:
This dataset contains 1647022 parallel speech-text pairs for Yoruba, a language spoken primarily in Nigeria and other West African countries. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Language: Yoruba - yo