Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Mbosi-TTS-Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
Ins
Host:
The dataset consists of paired audio and text data on Mbosi (mdw), a language spoken in Congo. The audio corpus consists of 2,575 clips read by one speaker totaling 275 min 48.35 sec. The dataset also contains a mapping file of audio and text with 2,597 lines. Each line begins with the name of an audio file, followed by a tab and then the corresponding text excerpt. This dataset is suitable for TTS tasks.

Visit

mozilladatacollective.com

Tasks

speech processingtext to speech

Languages

Mbosi

Tags

mdcmozilla data collectiveTTSWAVTSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similar

Kituba-TTS-DatasetLingala-TTS-DatasetTobydata Tts DatasetSuundi-TTS-DatasetKiswahili TTS DatasetRw Tts Dataset

Kituba-TTS-Dataset

Paired audio and text data on Kituba (mkw), a language spoken in Congo. The audio corpus consists of

Lingala-TTS-Dataset

The dataset contains audio and text resources in Lingala, a Bantu language spoken in the Republic of

Tobydata Tts Dataset

Luganda TTS dataset (Toby-data) collected by TericLab. Contains read speech in Luganda, primarily on

Suundi-TTS-Dataset

The dataset consists of paired audio and text data on Suundi (sdj), a language spoken in Congo. The

Kiswahili TTS Dataset

The dataset contains Kiswahili text and audio files. The dataset contains 7,108 text files and audio files. The Kiswahili dataset was created from an open-source non-copyrighted material: Kiswahili audio Bible. The authors permit use for non-profit, educational, a

Rw Tts Dataset

Kinyarwanda (rw) text-to-speech dataset. Studio-recorded read speech aligned with transcriptions, co