Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Nexdata-AI/300-Hour-Swahili-Tanzania-Speech-Dataset-Transcribed-Dialogue-Audio

Domain:

natural language processing

Record type:

dataset
Creator:
Nex
Host:
# 300-Hour-Swahili-Tanzania-Speech-Dataset-Transcribed-Dialogue-Audio --- language: - sw --- ## Description Swahili(Tanzania) Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied. For more details, please refer to the link: nexdata.ai ## Specifications ### Format 24kHz, 16 bit, wav, mono channel; ### Content category Dialogue based on given topics; ### Recording condition Low background noise (indoor); ### Recording device Android smartphone, iPhone; ### Speaker About 300 native speakers in total; ### Country Tanzania(TZA); ### Language(Region) Code sw-TZ; ### Language Swahili; ### Features of annotation Transcription text, timestamp, speaker ID, gender, noise,PII redacted. ### Accuracy Rate Word Accuracy Rate (WAR) 98% ## Licensing Information Commercial License

Visit

github.com

Tasks

automatic speech recognitionspeech processing

Languages

Swahili

Tags

asrspeech-recognitionspeech-to-texttext-to-speech

Similar

Swahili Audio Datasetshunyalabs/swahili-speech-dataset🇰🇪 Swahili Speech DatasetKallaama: A Transcribed Speech Dataset about Agriculture in the Three Most Widely Spoken Languages in Senegal Kallaama Wolof speech dataset Kallaama Pulaar speech dataset Kallaama Sereer speech dataset"Amharic Speech Corpus: A 20-Hour Multi-Speaker Dataset for Automatic Speech Recognition"Kallaama: A Transcribed Speech Dataset about Agriculture in Wolof, Pulaar, and Sereer

Swahili Audio Dataset

shunyalabs/swahili-speech-dataset

🇰🇪 Swahili Speech Dataset

🌍 Swahili — The lingua franca of East Africa. Spoken by 200+ million people across Kenya, Tanzania,

Kallaama: A Transcribed Speech Dataset about Agriculture in the Three Most Widely Spoken Languages in Senegal Kallaama Wolof speech dataset Kallaama Pulaar speech dataset Kallaama Sereer speech dataset

This data is transcribed speech data, in Wolof, Pulaar and Sereer. The recordings are about agricul

"Amharic Speech Corpus: A 20-Hour Multi-Speaker Dataset for Automatic Speech Recognition"

"This dataset is a 20.03-hour Amharic speech corpus recorded from 100 native speakers and containing

Kallaama: A Transcribed Speech Dataset about Agriculture in Wolof, Pulaar, and Sereer

Train Automatic Speech Recognition (ASR) models for Senegalese languages - Build Text-to-Speech (TTS