Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Large vocabulary speech recognition for languages of Africa: multilingual modeling and self-supervised learning

Domaine:

natural language processing

Type de record:

paper
Almost none of the 2,000+ languages spoken in Africa have widely available automatic speech recognition systems, and the required data is also only available for a few languages. We have experimented with two techniques which may provide pathways to large vocabulary speech recognition for African languages: multilingual modeling and self-supervised learning. We gathered available open source data and collected data for 15 languages, and trained experimental models using these techniques. Our results show that pooling the small amounts of data available in multilingual end-to-end models, and pretraining on unsupervised data can help improve speech recognition quality for many African languages.

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Languages

AkanHausaIgboKinyarwandaNdebeleSetswanaSotho, NorthernSotho, SouthernSwahiliSwati+5

Similaires

Multilingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitchingEnsemble of learning in Self-supervised speech recognitionLarge Vocabulary Spontaneous Speech Recognition for TigrignaAdvancing Arabic Speech Recognition Through Large-Scale Weakly Supervised LearningAn Amharic speech corpus for large vocabulary continuous speech recognitionMorph-based speech recognition and modeling of out-of-vocabulary words across languages

Multilingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitching

While many speakers of low-resource languages regularly code-switch between their languages and othe

Ensemble of learning in Self-supervised speech recognition

Ensemble of learning in Self-supervised speech recognition

Poster presented at the Deep Learning Indaba 2023 by Ussen Kimanuka

Large Vocabulary Spontaneous Speech Recognition for Tigrigna

This thesis proposes and describes a research attempt at designing and developing a speaker independ

Advancing Arabic Speech Recognition Through Large-Scale Weakly Supervised Learning

Automatic speech recognition (ASR) is crucial for human-machine interaction in diverse applications

An Amharic speech corpus for large vocabulary continuous speech recognition

Morph-based speech recognition and modeling of out-of-vocabulary words across languages

We explore the use of morph-based language models in large-vocabulary continuous-speech recognition