Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Using Radio Archives for Low-Resource Speech Recognition: Towards an Intelligent Virtual Assistant for Illiterate Users

Domaine:

natural language processing

Type de record:

paper
For many of the 700 million illiterate people around the world, speech recognition technology could provide a bridge to valuable information and services. Yet, those most in need of this technology are often the most underserved by it. In many countries, illiterate people tend to speak only low-resource languages, for which the datasets necessary for speech technology development are scarce. In this paper, we investigate the effectiveness of unsupervised speech representation learning on noisy radio broadcasting archives, which are abundant even in low-resource languages. We make three core contributions. First, we release two datasets to the research community. The first, West African Radio Corpus, contains 142 hours of audio in more than 10 languages with a labeled validation subset. The second, West African Virtual Assistant Speech Recognition Corpus, consists of 10K labeled audio clips in four languages. Next, we share West African wav2vec, a speech encoder trained on the noisy radio corpus, and compare it with the baseline Facebook speech encoder trained on six times more data of higher quality. We show that West African wav2vec performs similarly to the baseline on a multilingual speech recognition task, and significantly outperforms the baseline on a West African language identification task. Finally, we share the first-ever speech recognition models for Maninka, Pular and Susu, languages spoken by a combined 10 million people in over seven countries, including six where the majority of the adult population is illiterate. Our contributions offer a path forward for ethical AI research to serve the needs of those most disadvantaged by the digital divide.

Visit

arxiv.orgwww.aaai.org

Connected records

datasetdatasetproject

Tasks

automatic speech recognitiontext to speechlanguage identificationspeech processing

Languages

Kisi, SouthernKonoManinka, KonyankaManinkakan, EasternManinkakan, WesternPularSusuTomaXaasongaxango

Tags

speech representation learning

Licenses

Creative Commons Attribution-ShareAlike 4.0 International License

Similaires

West African Virtual Assistant Speech Recognition CorpusDevelopment of an intelligent virtual assistant for digitalization of Moroccan agricultureRobust speech recognition for low-resource languagesImproved Meta Learning for Low Resource Speech RecognitionDEFI-COLaF/Speech-Recognition-for-Low-Resource-LanguagesText-To-Speech Data Augmentation for Low Resource Speech Recognition

West African Virtual Assistant Speech Recognition Corpus

This dataset contains 10,083 recorded utterances in French, Maninka, Pular and Susu from 49 speakers (16 female and 33 male) ranging from 5 to 76 years old on a variety of devices. Please see our paper for more details on this dataset. Additional resources can be

Development of an intelligent virtual assistant for digitalization of Moroccan agriculture

This paper presents the design, development, and implementation of an innovative text-to-text chatbo

Robust speech recognition for low-resource languages

Process of human-machine interaction is an integral part of everyday human life in a modern world. T

Improved Meta Learning for Low Resource Speech Recognition

We propose a new meta learning based framework for low resource speech recognition that improves the

DEFI-COLaF/Speech-Recognition-for-Low-Resource-Languages

# Speech-Recognition-for-Low-Resource-Languages This repository contains code to fine-tune Whisper

Text-To-Speech Data Augmentation for Low Resource Speech Recognition

Nowadays, the main problem of deep learning techniques used in the development of automatic speech r