Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

stano03/jambogpt-massive-dataset

Domain:

natural language processing

Record type:

dataset
Creator:
sta
Host:
This is the world's largest open-source African language voice dataset with 100,000 hours of high-quality audio in 10 African languages. Metric Value Total Hours 100,000 Total Samples 5,000,000+ Total Speakers 10,000+ Languages 10 Audio Quality 16kHz, 16-bit PCM License CC-BY-4.0 East Africa (30,000 hours)

Visit

huggingface.co

Tasks

speech processing

Similar

stano03/jambogpt-real-datasetstano03/jambogpt-kenyan-languagesstano03/jambogpt-swahili-tts-v1

stano03/jambogpt-real-dataset

Real audio samples for 10 African languages, ready for training speech recognition and text-to-speec

stano03/jambogpt-kenyan-languages

This is a comprehensive voice dataset for Kenyan languages, created to advance AI accessibility in A

stano03/jambogpt-swahili-tts-v1