Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

DhoNam: Dholuo Speech dataset

Domain:

natural language processing

Record type:

dataset
Creator:
Mas
Host:
DhoNam: Dholuo Speech dataset is a speech corpus designed to supercharge Automatic Speech Recognition (ASR) and other speech technologies for Dholuo, one of Kenya’s major indigenous languages. This dataset contains native-speaker audio recordings collected through a platform where users read aloud a displayed sentence. The dataset includes the audio recordings and the corresponding prompt/sentence that was read.

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionspeech processing

Languages

Dholuo

Tags

mdcmozilla data collectiveASRWEBM

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)