Logo Lanfrica

Making AI understand 3 East African languages: Kiswahili, Kinyarwanda and Luganda - Open-source speech-to-text datasets - Mozilla Common Voice

Domaine:

natural language processing

Type de record:

dataset
By collecting more than 1064 hours of AI recorded speech in Kiswahili (by 03/2026), this effort created the largest open-source voice dataset of diverse Swahili speakers for speech recognition (speech-to-text). You can use this resource to build AI systems understanding spoken… Notes / challenges: Fair Forward portfolio. Dataset

Similaires