Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Beembe-TTS-Dataset

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Ins
Hôte:
The dataset consists of paired audio and text data on Beembe (beq), a language spoken in Congo. The audio corpus consists of 6,933 clips read by one speaker totaling 275 min 48.35 sec. The dataset also contains a mapping file of audio and text with 4,422 lines. Each line begins with the name of an audio file, followed by a tab and then the corresponding text excerpt. This dataset is suitable for TTS tasks.

Visit

mozilladatacollective.com

Tasks

speech processingtext to speech

Languages

Beembe

Tags

mdcmozilla data collectiveTTSWAVTSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similaires

Kituba-TTS-DatasetLingala-TTS-DatasetTobydata Tts DatasetSuundi-TTS-DatasetKiswahili TTS DatasetRw Tts Dataset

Kituba-TTS-Dataset

Paired audio and text data on Kituba (mkw), a language spoken in Congo. The audio corpus consists of

Lingala-TTS-Dataset

The dataset contains audio and text resources in Lingala, a Bantu language spoken in the Republic of

Tobydata Tts Dataset

Luganda TTS dataset (Toby-data) collected by TericLab. Contains read speech in Luganda, primarily on

Suundi-TTS-Dataset

The dataset consists of paired audio and text data on Suundi (sdj), a language spoken in Congo. The

Kiswahili TTS Dataset

The dataset contains Kiswahili text and audio files. The dataset contains 7,108 text files and audio files. The Kiswahili dataset was created from an open-source non-copyrighted material: Kiswahili audio Bible. The authors permit use for non-profit, educational, a

Rw Tts Dataset

Kinyarwanda (rw) text-to-speech dataset. Studio-recorded read speech aligned with transcriptions, co