Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

stano03/jambogpt-real-dataset

Domain:

natural language processing

Record type:

dataset
Creator:
sta
Host:
Real audio samples for 10 African languages, ready for training speech recognition and text-to-speech models. Total Samples: 500 Languages: 10 African languages Audio Quality: 16kHz, 16-bit PCM WAV License: CC-BY-4.0 (Open Source) 🇰🇪 Swahili (50 samples) 🇰🇪 Kikuyu (50 samples) 🇳🇬 Yoruba (50 samples) 🇳🇬 Hausa (50 samples) 🇪🇹 Amharic (50 samples)

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processingtext to speech

Languages

AmharicGikuyuHausaSwahiliYoruba

Similar

stano03/jambogpt-massive-datasetstano03/jambogpt-kenyan-languagesstano03/jambogpt-swahili-tts-v1

stano03/jambogpt-massive-dataset

This is the world's largest open-source African language voice dataset with 100,000 hours of high-qu

stano03/jambogpt-kenyan-languages

This is a comprehensive voice dataset for Kenyan languages, created to advance AI accessibility in A

stano03/jambogpt-swahili-tts-v1