Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Kituba-TTS-Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
Ins
Host:
Paired audio and text data on Kituba (mkw), a language spoken in Congo. The audio corpus consists of 8,302 clips read by one speaker, totalling 350 min 11.98 sec. The dataset also contains a mapping file of audio and text with 8,173 lines. Each line begins with the name of an audio file, followed by a tab and then the corresponding text excerpt. This dataset is suitable for TTS tasks.

Visit

mozilladatacollective.com

Tasks

speech processingtext to speech

Languages

KitubaKituba

Tags

mdcmozilla data collectiveTTSWAVTSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similar

Laari-TTS-DatasetLingala-TTS-DatasetTobydata Tts DatasetMbosi-TTS-DatasetKiswahili TTS DatasetRw Tts Dataset

Laari-TTS-Dataset

The dataset contains audio and text resources on Laari, a Bantu language spoken in the Congo. The re

Lingala-TTS-Dataset

The dataset contains audio and text resources in Lingala, a Bantu language spoken in the Republic of

Tobydata Tts Dataset

Luganda TTS dataset (Toby-data) collected by TericLab. Contains read speech in Luganda, primarily on

Mbosi-TTS-Dataset

The dataset consists of paired audio and text data on Mbosi (mdw), a language spoken in Congo. The a

Kiswahili TTS Dataset

The dataset contains Kiswahili text and audio files. The dataset contains 7,108 text files and audio files. The Kiswahili dataset was created from an open-source non-copyrighted material: Kiswahili audio Bible. The authors permit use for non-profit, educational, a

Rw Tts Dataset

Kinyarwanda (rw) text-to-speech dataset. Studio-recorded read speech aligned with transcriptions, co