Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Fine-Tuning DeepSpeech Speech-To-Text Model for Nigerian English and Yoruba-English Code-Switched Speech

Domaine:

natural language processing

Type de record:

papermodeldataset
Créateur:
OloOloAda
Éditeur:
DepDep
Éditeur:
CCSD
Hôte:avatar
International audience Speech-to-Text (STT) systems, despite their stellar performance in recent years, still struggle with recognising non-Western English accents and speech that features Code-Switching (CS), a linguistic phenomenon common in regions such as Nigeria. This study addresses that challenge for Nigerian English and Yoruba-English code-switched speech by adapting Mozilla’s DeepSpeech 0.9.3 model and fine-tuning it using a custom dataset of 118 minutes (approximately 1.97 hours). This process involved transfer learning and hyperparameter optimisation over iterative training sessions on a CPU-based setup. The model’s performance was evaluated using Word Error Rate (WER) and Character Error Rate (CER), with the best model showing modest improvements over the baseline model and achieving a WER of 0.760261 and CER of 0.381241 after 55 epochs. Although limited computing resources and the small dataset imposed significant constraints on the work, the study demonstrated the potential of fine-tuning and transfer learning for model adaptation to low-resource languages and code-switching contexts. Future work will require access to GPU resources for improved convergence and transcription accuracy, an expanded dataset and support for Yoruba diacritics to improve the quality of transcriptions.

Visit

hal.science

Tasks

automatic speech recognitioncode switchingspeech processing

Languages

Yoruba

Tags

[INFO]Computer Science [cs]

Similaires

Low-Resource Hate Speech Detection in English-Swahili Code-Switched Text Using Fine-Tuning of Pre-trained Language ModelsChichewa-English Code-Switched Speech DatasetCorrection to: Two sepedi‑english code‑switched speech corporaEnglish-IsiZulu Code-Switched Speech Recognition DatasetGhana English-Twi Code-switched Speech CorpusAn analysis of fine-tuned representations for code-switched speech recognition of Yorùbá and English

Low-Resource Hate Speech Detection in English-Swahili Code-Switched Text Using Fine-Tuning of Pre-trained Language Models

The use of social media in East Africa has grown rapidly, and with it, the spread of hate speech has

Chichewa-English Code-Switched Speech Dataset

A speech dataset containing 247 audio recordings of Chichewa-English code-switched phrases. Code-swi

Correction to: Two sepedi‑english code‑switched speech corpora

English-IsiZulu Code-Switched Speech Recognition Dataset

Dataset for semi-supervised acoustic and language model training for English-isiZulu code-switched s

Ghana English-Twi Code-switched Speech Corpus

Gold-standard English-Twi code-switched speech with transcripts and speaker meta

An analysis of fine-tuned representations for code-switched speech recognition of Yorùbá and English

An analysis of fine-tuned representations for code-switched speech recognition of Yorùbá and English

Poster presented at the Deep Learning Indaba 2022 by Tolúlọpẹ́ Ògúnrẹ̀mí