Logo Lanfrica

Rafat-decodis/Robust-ASR-for-Low-Resource-Languages

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Raf
Hôte:
Exploring Benchmark Gaps and Real-World Speech Generalization for Language in Low Resource # 🧠 A Robust ASR Study for Swahili and Yoruba ## 🔍 Exploring Benchmark Gaps and Real-World Speech Generalization Welcome to our open-source repository exploring **automatic speech recognition (ASR)** in **Swahili** and **Yoruba**, where we investigate how open benchmark datasets compare to **real-world, noisy, conversational speech**. We present a complete ablation study using: - 🌐 Open-source datasets - 🗣️ A hand-collected, high-variance real-world dataset - 🧪 Three ASR models: Open-Source-only, Custom-only, Combined - Attached presentation for the ablation --- ## 🚀 Why This Matters Low-resource African languages like Swahili and Yoruba are often trained and evaluated on clean, controlled datasets. But real-world usage involves: - Noisy environments (phones, streets, homes) - Spontaneous, conversational phrasing (Non-read out) - Speaker and device diversity - Non-standard syntax and informal speech - Groups not-involved in data freelance 🔎 We show that **benchmark performance is not a reliable proxy for real-world ASR quality** — and that even a small, well-labeled domain dataset can **improve generalization** meaningfully. Decodis collects “in the wild” natural language data sets in the course of conducting research with populations who speak low-resource languages. The audio data is collected by Interactive Voice Recording (IVR) by calling a basic or smartphone. The interview questions prompt interviewees to answer in open-ended responses. One interview can be completed by 1000+ people in one week, quickly yielding hundreds of hours of natural language data. --- ## 📦 Datasets Used **Swahili** | Dataset | Type | Langs | Hours | Notes | |--------|------|-------|--------|-------| | **Open Source Dataset** | Converstional, Read-aloud, crowd-sourced | Swahili | ~400h |Clean studio-style and noisy speech, Mixed mic quality, many speakers, manually labelled | | **Decodis Dataset** | Conversational, noisy, real-world | Swahili | ~50h | Diverse, high-noise, …