Moonshine STT for African languages — Msaidizi
# Moonshine African Languages
**Building production-grade African language speech recognition for Msaidizi — at $0 cost.**
## What Is This?
This project fine-tunes Moonshine AI (open-source STT) for 6 African languages spoken by Msaidizi's informal worker users.
Msaidizi is Africa-first. Current models (Whisper, Google STT) are terrible at African dialects. Moonshine's architecture lets us build **adaptive, dialect-aware models** that learn from Msaidizi's users — something Whisper can never do.
## Datasets (6,500+ Hours FREE)
The dataset landscape transformed in 2026. Every language Msaidizi needs now has massive free data:
| Dataset | Hours | Languages | Source |
|---------|-------|-----------|--------|
| **AfriVoices-KE** (Maseno Univ, Apr 2026) | 3,000 | Dholuo, Kikuyu, Kalenjin, Maasai, Somali | CC BY 4.0 |
| **Thiomi Dataset** (Harvard, Apr 2026) | 1,500 | Swahili, Kikuyu, Kamba, Luo, Kipsigis + 5 more | CC BY 4.0 |
| **Google WAXAL** (Feb 2026) | 1,485 | Swahili, Luo, Kikuyu + 21 more | CC BY 4.0 |
| **Mozilla Common Voice** | 400+ | Swahili (200h+), Luo (120h), Kalenjin (92h) | CC0 |
| **Kencorpus** (2023) | 177 | Swahili, Dholuo, Luhya | Open |
**Every language has data. No more "collect first" — train everything now.**
## What Is This?
This project fine-tunes Moonshine AI (open-source STT) for 6 African languages spoken by Msaidizi's informal worker users:
| Language | Status | Mode | Data Source |
|----------|--------|------|------------|
| Swahili | 🟢 Ready to train | Full transcription | Common Voice + ALFFA + Thiomi + WAXAL (~600h+) |
| Luo | 🟢 Ready to train | Full transcription | Common Voice (120h) + AfriVoices-KE + Thiomi + WAXAL |
| Kikuyu | 🟢 Ready to train | Full transcription | AfriVoices-KE + Thiomi + WAXAL |
| Kamba | 🟡 Has data | Command recognition | Thiomi Dataset (Harvard 2026) |
| Luhya | 🟡 Has data | Command recognition | Kencorpus (177h) |
| Kalenjin | 🟢 Ready to train | Command recognition | Common Voice (92h) + AfriVoices- …