Logo Lanfrica

Niger Volta LTI: Yoruba Audio

Domaine:

natural language processing

Type de record:

dataset
This repo aggregates audio/speech corpora for Yorùbá tasks. The corpora may contain aligned text or be purely unlabeled. The objective is to have a bird's eye view of available Yorùbá audio, and it's metadata and entropy, to inform additional data collection tasks & modeling. For example, if we see a large Broadcast news corpus, we might be interested to train a self-supervised model on a pretext task to generate speech embeddings for use in ASR/TTS work.

Languages

Licenses