Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

🇰🇪 Swahili Speech Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
Sil
Host:
🌍 Swahili — The lingua franca of East Africa. Spoken by 200+ million people across Kenya, Tanzania, Uganda, Rwanda, Burundi, and the DRC. 📧 Need more? sofia@silencioai.com — we have 9,786 hours of Swahili voice data. 47 high-quality Swahili recordings (~21 minutes) from native speakers across East Africa. Language Speakers Regions Sample Size 🇰🇪 Kiswahili Native speakers Kenya, Tanzania, Uganda

Visit

huggingface.co

Tasks

speech processing

Languages

SwahiliSwahili, CoastalSwahili, Congo

Tags

swahilikiswahilieast-africakenyatanzaniaugandarwandaburundidrcafrican-languages+6

Licenses

cc-by-nc-4.0

Similar

shunyalabs/swahili-speech-datasetSwahili Words Speech-Text Parallel DatasetSWAHILI AND CODE-SWITCHED ENGLISH-SWAHILI POLITICAL HATE SPEECH DETECTION TEXTUAL DATASETA Curated Crowdsourced Dataset of Luganda and Swahili Speech for Text-to-Speech SynthesisNexdata-AI/300-Hour-Swahili-Tanzania-Speech-Dataset-Transcribed-Dialogue-AudioSwahili Speech Dataset Development and Improved Pre-training Method for Spoken Digit Recognition

shunyalabs/swahili-speech-dataset

Swahili Words Speech-Text Parallel Dataset

This dataset contains 411048 parallel speech-text pairs for Swahili, a widely spoken language in Eas

SWAHILI AND CODE-SWITCHED ENGLISH-SWAHILI POLITICAL HATE SPEECH DETECTION TEXTUAL DATASET

This dataset consist of swahili and code switched English-Swahili tweets labeled for hate speech typ

A Curated Crowdsourced Dataset of Luganda and Swahili Speech for Text-to-Speech Synthesis

This dataset contains curated and preprocessed speech recordings in Luganda and Kiswahili for use in

Nexdata-AI/300-Hour-Swahili-Tanzania-Speech-Dataset-Transcribed-Dialogue-Audio

# 300-Hour-Swahili-Tanzania-Speech-Dataset-Transcribed-Dialogue-Audio --- language: - sw --- ## De

Swahili Speech Dataset Development and Improved Pre-training Method for Spoken Digit Recognition

Speech dataset is an essential component in building commercial speech applications. However, low-re