Logo Lanfrica

Kachiengineers/Ehugbo-Audio

Domain:

natural language processing

Record type:

dataset
Creator:
Kac
Host:
This is a repository of my Ehugbo project working on the first publicly available Ehugbo audio data sets # Ehugbo ASR: Benchmarking Speech Recognition for a Low-Resource Igbo Dialect This repository contains the code and resources for the paper *"When Endangered Voices Speak: Building the First Ehugbo Dialect Audio Dataset Through Grassroots Collaboration"*, presented at the WiML 2025 Workshop @ NeurIPS. Mozilla Data Collective for the full dataset: datacollective.mozillafound… ## Abstract The rapid advancement of speech technologies has created a growing linguistic divide, leaving behind communities whose languages are underrepresented. This is especially true for endangered languages like **Ehugbo**, a dialect of Igbo spoken by approximately 150,000 people in Afikpo, Nigeria. With no prior publicly available audio data, Ehugbo is at high risk of digital extinction. This project details the collaborative, grassroots construction of the first open-source Ehugbo audio dataset—a 42.99-minute corpus of biblical recitations. We benchmark several pre-trained ASR models on this new dataset, establishing baseline performance and highlighting the unique challenges of dialectal, low-resource ASR. ## The Ehugbo ASR Dataset Our work introduces the first publicly available audio dataset for the Ehugbo dialect. * **Language:** Ehugbo (a dialect of Igbo) * **Total Duration:** 42.99 minutes * **Content Source:** The Ehugbo New Testament (2020 Publication) * **Data Collection:** The project was built from the ground up in collaboration with the local community, involving six native speakers for recording and three validators for quality assurance. * **Availability:** The dataset is still being curated and to be released publicly soon, however, to reproduce this experiment, feel free to reach out. ## Benchmark Results We evaluated several existing Igbo ASR models on our dataset to establish a performance baseline. The `CLEAR-Global/w2v-bert-2.0-igbo_naijavoices_250h` model, fine-tuned on a large corpus of diverse N …