Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Read Speech in Kenyan Swahili (6h)

Domaine:

natural language processing

Type de record:

dataset
Créateur:
CLEAR Global
Hôte:
A single-speaker read speech dataset in Kenyan Swahili, produced as part of CLEAR Global's Gamayun Language Data Kits initiative. The dataset contains 4,700 pre-segmented utterances (~6 hours, 21,852 seconds) recorded by an anonymous male Kenyan speaker. Sentences were prompted from a script: Swahili translations of general-domain English sentences sourced from the Tatoeba repository. The same sentence set is used in CLEAR Global's Gamayun Swahili–English parallel text kit. The archive includes WAV audio files organised by recording session and a metadata TSV with transcriptions, file paths, and durations.

Visit

mozilladatacollective.com

Tasks

speech processing

Languages

Swahili

Tags

mdcmozilla data collectiveASRWAVTSV

Licenses

Creative Commons Attribution Non Commercial 4.0 International (CC-BY-NC-4.0)

Similaires

Read Speech in Kenyan Swahili (6h)Read Speech in Kenyan Swahili (6h)

Read Speech in Kenyan Swahili (6h)

A single-speaker read speech dataset in Kenyan Swahili, containing approximately 6 hours of prompted

Read Speech in Kenyan Swahili (6h)

A single-speaker read speech dataset in Kenyan Swahili, containing approximately 6 hours of prompted