Exploring Benchmark Gaps and Real-World Speech Generalization for Language in Low Resource
# 🧠 A Robust ASR Study for Swahili and Yoruba
## 🔍 Exploring Benchmark Gaps and Real-World Speech Generalization
Welcome to our open-source repository exploring **automatic speech recognition (ASR)** in **Swahili** and **Yoruba**, where we investigate how open benchmark datasets compare to **real-world, noisy, conversational speech**.
We present a complete ablation study using:
- 🌐 Open-source datasets
- 🗣️ A hand-collected, high-variance real-world dataset
- 🧪 Three ASR models: Open-Source-only, Custom-only, Combined
- Attached presentation for the ablation
---
## 🚀 Why This Matters
Low-resource African languages like Swahili and Yoruba are often trained and evaluated on clean, controlled datasets. But real-world usage involves:
- Noisy environments (phones, streets, homes)
- Spontaneous, conversational phrasing (Non-read out)
- Speaker and device diversity
- Non-standard syntax and informal speech
- Groups not-involved in data freelance
🔎 We show that **benchmark performance is not a reliable proxy for real-world ASR quality** — and that even a small, well-labeled domain dataset can **improve generalization** meaningfully.
Decodis collects “in the wild” natural language data sets in the course of conducting research with populations who speak low-resource languages. The audio data is collected by Interactive Voice Recording (IVR) by calling a basic or smartphone. The interview questions prompt interviewees to answer in open-ended responses. One interview can be completed by 1000+ people in one week, quickly yielding hundreds of hours of natural language data.
---
## 📦 Datasets Used
**Swahili**
| Dataset | Type | Langs | Hours | Notes |
|--------|------|-------|--------|-------|
| **Open Source Dataset** | Converstional, Read-aloud, crowd-sourced | Swahili | ~400h |Clean studio-style and noisy speech, Mixed mic quality, many speakers, manually labelled |
| **Decodis Dataset** | Conversational, noisy, real-world | Swahili | ~50h | Diverse, high-noise, …