# Towards End-to-End Training of Automatic Speech Recognition for Nigerian Pidgin
## Overview
This project explores **end-to-end training of Automatic Speech Recognition (ASR) systems for Nigerian Pidgin**, a widely creole spoken language across West Africa but underrepresented in speech technologies.
Despite being spoken by over **75 million people**, the Pidgin language lacks sufficient digital resources and language technologies. This work addresses the gap by building datasets and training ASR models tailored for a variant of the pidgin language - Nigerian Pidgin, contributing towards more inclusive and equitable language technologies.
- **Project Website:** ASR Nigerian Pidgin
- **Interactive Demo:** Hugging Face Space
- **Dataset:** Hugging Face Datasets
- **Paper:** Arvix
---
>**Towards End-to-End Training of Automatic Speech Recognition for Nigerian Pidgin**\
>Amina Mardiyyah Rufai, Abeeb Afolabi, Daniel Ajisafe, Oluwabukola Adegboro, Esther Oduntan, and Tayo Arulogun
----
## Problem
Most state-of-the-art ASR systems are trained on high-resource languages (e.g., English, German, Japanese). This creates a significant gap for low-resource languages like Nigerian Pidgin, where lack of data and research limits the development of speech-enabled applications.
**Goal:** This project focuses on the development of an end-to-end speech recognition system customised for Nigerian Pidgin English.
---
## Approach
We implemented and compared three(3) end-to-end ASR training methods, experimenting with different ASR pipelines and architectures:
1. **Wav2Vec2-Base**
2. **Wav2Vec2-XLSR53**
3. **NVIDIA NeMo**
Each model was trained and evaluated on our Nigerian Pidgin speech corpus. Our final and best performing model on the test set was the **Wav2Vec2-XLSR53**, achieving a reduced error rate of 29.6%.
## Key Contributions
- A publicly accessible ASR system for Nigerian Pidgin
- Free speech corpus for Nigerian Pidgin
- First parallel (speech-to-text) Nigerian …