Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Read Speech in Kenyan Swahili (6h)

Domain:

natural language processing

Record type:

dataset
Creator:
CLEAR Global
Host:
A single-speaker read speech dataset in Kenyan Swahili, produced as part of CLEAR Global's Gamayun Language Data Kits initiative. The dataset contains 4,700 pre-segmented utterances (~6 hours, 21,852 seconds) recorded by an anonymous male Kenyan speaker. Sentences were prompted from a script: Swahili translations of general-domain English sentences sourced from the Tatoeba repository. The same sentence set is used in CLEAR Global's Gamayun Swahili–English parallel text kit. The archive includes WAV audio files organised by recording session and a metadata TSV with transcriptions, file paths, and durations.

Visit

mozilladatacollective.com

Tasks

speech processing

Languages

Swahili

Tags

mdcmozilla data collectiveASRWAVTSV

Licenses

Creative Commons Attribution Non Commercial 4.0 International (CC-BY-NC-4.0)

Similar

Read Speech in Kenyan Swahili (6h)Read Speech in Kenyan Swahili (6h)

Read Speech in Kenyan Swahili (6h)

A single-speaker read speech dataset in Kenyan Swahili, containing approximately 6 hours of prompted

Read Speech in Kenyan Swahili (6h)

A single-speaker read speech dataset in Kenyan Swahili, containing approximately 6 hours of prompted