# ASP_yoruba
We are solving a speech recognition problem for the Yoruba language.
**As promising models for training , the following were considered:**
1. facebook/wav2vec2-xls-r-300m
2. openai/whisper-small
3. Ayoola/cdial-yoruba-test (Ayoola/cdial-yoruba-test)
**We use datasets for training**:
1. google/fleurs (FLEURS: Few-shot Learning E…)
2. bibleTTS (Bible TTS)
3. mozilla-foundation/common_voice_12_0
4. Lagos-NWU (repo.sadilar.org)
WER (Word error rate) is used as the main evaluation metric.
**WER**:
1. *67.8866*
**Dataset**: google/fleurs
**Model**: openai/whisper-small
**Source**: steja/whisper-small-yoruba
2. *57.8640*
**Dataset**: google/fleurs + mozilla-foundation/common_voice_12_0
**Model**: facebook/wav2vec2-xls-r-300m
**Source**: huggingface.co