Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Common Voice 17.0 (Swahili, Kenyan Sample)

Domain:

natural language processing

Record type:

dataset
Creator:
Ver
Host:
This dataset is a filtered sample of the Mozilla Common Voice 17.0 corpus, focusing on Swahili (sw) speech with Kenyan voices. It has been subsetted for experimentation and prototyping in ASR (Automatic Speech Recognition) models targeting speech-impaired users in Kenya, covering Kenyan English and Kiswahili. Language: Kiswahili (Swahili, sw) Accent/Region: Kenyan speakers

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Languages

SwahiliSwahili, CoastalSwahili, Congo

Tags

swahilispeechkenyanaudiocommon-voiceautomatic-speech-recognition

Licenses

cc0-1.0

Similar

Common Voice Corpus 17.0EYEDOL/swahili-common-voice-woman_soundCommon Voice Scripted Speech 26.0 - SwahiliBenjamin-png/swahili-common-voice-woman_soundCommon Voice

Common Voice Corpus 17.0

This dataset is an unofficial version of the Mozilla Common Voice Corpus 17. It was downloaded and c

EYEDOL/swahili-common-voice-woman_sound

Common Voice Scripted Speech 26.0 - Swahili

A collection of read speech recordings in Swahili (Kiswahili).

Benjamin-png/swahili-common-voice-woman_sound

Common Voice

Common Voice is Mozilla's initiative to help teach machines how real people speak. The dataset currently consists of 7,335 validated hours of speech in 60 languages, but we’re always adding more voices and languages.