Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

ddd-kenya-somali-68hrs-asr-part1

Domain:

natural language processing

Record type:

dataset
Creator:
Dig
Host:
This dataset, curated by Digital Divide Data (DDD), provides high-quality audio recordings and corresponding text transcriptions for the Somali (som) language. The collection includes thousands of unique utterances per language to support diverse acoustic modeling. All transcriptions have undergone a manual verification process to ensure high linguistic accuracy. Recordings feature a balanced mix of genders and various age groups to minimize bias in downstream AI models. This data is specifically designed for training Automatic Speech Recognition (ASR) systems, Text-to-Speech (TTS) synthesis, and general linguistic research for underrepresented African languages.

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionspeech processingtext to speech

Languages

Somali

Tags

mdcmozilla data collectiveASRWAVTSV

Licenses

Creative Commons Attribution 4.0 International (CC-BY-4.0)