Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

A kiswahili Dataset for Development of Text-To-Speech System

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Kip
Éditeur:
Kip
Éditeur:
Men
Hôte:avatar
The dataset contains Kiswahili text and audio files. The dataset contains 7,108 text files and audio files. The Kiswahili dataset was created from an open-source non-copyrighted material: Kiswahili audio Bible. The authors permit use for non-profit, educational, and public benefit purposes. The downloaded audio files length was more than 12.5s. Therefore, the audio files were programmatically split into short audio clips based on silence. They were then combined based on a random length such that each eventual audio file lies between 1 to 12.5s. This was done using python 3. The audio files were saved as a single channel,16 PCM WAVE file with a sampling rate of 22.05 kHz The dataset contains approximately 106,000 Kiswahili words. The words were then transcribed into mean words of 14.96 per text file and saved in CSV format. Each text file was divided into three parts: unique ID, transcribed words, and normalized words. A unique ID is a number assigned to each text file. The transcribed words are the text spoken by a reader. Normalized texts are the expansion of abbreviations and numbers into full words. An audio file split was assigned a unique ID, the same as the text file.

Visit

doi.orgdata.mendeley.com

Tasks

text to speechspeech processing

Languages

SwahiliSwahili, CoastalSwahili, Congo

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Development of a Kiswahili text to speech systemDevelopment of a Kiswahili Text-to-Speech System Based on Tacotron 2 and WaveNet VocoderDevelopment of a Kiswahili Text-to-Speech System based on Tacotron 2 and Wave Net VocoderNeuro-growth/kiswahili-kikuyu-text-to-speech-Development of an Amharic text-to-speech system using cepstral methodA Curated Crowdsourced Dataset of Luganda and Swahili Speech for Text-to-Speech Synthesis

Development of a Kiswahili text to speech system

Development of a Kiswahili Text-to-Speech System Based on Tacotron 2 and WaveNet Vocoder

Development of a Kiswahili Text-to-Speech System based on Tacotron 2 and Wave Net Vocoder

Neuro-growth/kiswahili-kikuyu-text-to-speech-

bug free # Kikuyu Text-to-Speech App A Next.js application that translates English or Kiswahili te

Development of an Amharic text-to-speech system using cepstral method

A Curated Crowdsourced Dataset of Luganda and Swahili Speech for Text-to-Speech Synthesis

This dataset contains curated and preprocessed speech recordings in Luganda and Kiswahili for use in