Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Nexdata-AI/300-Hour-Swahili-Tanzania-Speech-Dataset-Transcribed-Dialogue-Audio

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Nex
Hôte:
# 300-Hour-Swahili-Tanzania-Speech-Dataset-Transcribed-Dialogue-Audio --- language: - sw --- ## Description Swahili(Tanzania) Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied. For more details, please refer to the link: nexdata.ai ## Specifications ### Format 24kHz, 16 bit, wav, mono channel; ### Content category Dialogue based on given topics; ### Recording condition Low background noise (indoor); ### Recording device Android smartphone, iPhone; ### Speaker About 300 native speakers in total; ### Country Tanzania(TZA); ### Language(Region) Code sw-TZ; ### Language Swahili; ### Features of annotation Transcription text, timestamp, speaker ID, gender, noise,PII redacted. ### Accuracy Rate Word Accuracy Rate (WAR) 98% ## Licensing Information Commercial License

Visit

github.com

Tasks

automatic speech recognitionspeech processing

Languages

Swahili

Tags

asrspeech-recognitionspeech-to-texttext-to-speech

Similaires

Swahili Audio Datasetshunyalabs/swahili-speech-dataset🇰🇪 Swahili Speech DatasetKallaama: A Transcribed Speech Dataset about Agriculture in the Three Most Widely Spoken Languages in Senegal Kallaama Wolof speech dataset Kallaama Pulaar speech dataset Kallaama Sereer speech dataset"Amharic Speech Corpus: A 20-Hour Multi-Speaker Dataset for Automatic Speech Recognition"Kallaama: A Transcribed Speech Dataset about Agriculture in Wolof, Pulaar, and Sereer

Swahili Audio Dataset

shunyalabs/swahili-speech-dataset

🇰🇪 Swahili Speech Dataset

🌍 Swahili — The lingua franca of East Africa. Spoken by 200+ million people across Kenya, Tanzania,

Kallaama: A Transcribed Speech Dataset about Agriculture in the Three Most Widely Spoken Languages in Senegal Kallaama Wolof speech dataset Kallaama Pulaar speech dataset Kallaama Sereer speech dataset

This data is transcribed speech data, in Wolof, Pulaar and Sereer. The recordings are about agricul

"Amharic Speech Corpus: A 20-Hour Multi-Speaker Dataset for Automatic Speech Recognition"

"This dataset is a 20.03-hour Amharic speech corpus recorded from 100 native speakers and containing

Kallaama: A Transcribed Speech Dataset about Agriculture in Wolof, Pulaar, and Sereer

Train Automatic Speech Recognition (ASR) models for Senegalese languages - Build Text-to-Speech (TTS