Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Fast Development of ASR in African Languages using Self Supervised Speech Representation Learning

Domaine:

natural language processing

Type de record:

project
Créateur:
MohThoNdoBes
Éditeur:
arXiv
Hôte:avatar
This paper describes the results of an informal collaboration launched during the African Master of Machine Intelligence (AMMI) in June 2020. After a series of lectures and labs on speech data collection using mobile applications and on self-supervised representation learning from speech, a small group of students and the lecturer continued working on automatic speech recognition (ASR) project for three languages: Wolof, Ga, and Somali. This paper describes how data was collected and ASR systems developed with a small amount (1h) of transcribed speech as training data. In these low resource conditions, pre-training a model on large amounts of raw speech was fundamental for the efficiency of ASR systems developed. Accepted at AfricaNLP2021 workshop at EACL 2021

Visit

doi.orgarxiv.org

Tasks

automatic speech recognitionspeech processing

Languages

SomaliWolof

Tags

Sound (cs.SD)Computation and Language (cs.CL)Audio and Speech Processing (eess.AS)FOS: Computer and information sciencesFOS: Computer and information sciencesFOS: Electrical engineering, electronic engineering, information engineeringFOS: Electrical engineering, electronic engineering, information engineering

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode