Logo Lanfrica

Yaka-TTS-Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
Ins
Host:
Paired audio and text data on Yaka (also known as West Teke), a language spoken in Congo. The audio corpus consists of 7,648 clips read by one speaker for a total duration of 344 min 40.48 sec. The dataset also contains a mapping file of audio and text with 7,648 lines. Each line begins with the name of an audio file, followed by a tab and then the corresponding text excerpt. This dataset is suitable for TTS tasks.