Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

West African Radio Corpus

Domaine:

natural language processing

Type de record:

dataset
This dataset contains 17,090 audio clips of length 30 seconds sampled from archives collected from 6 Guinean radio stations. The broadcasts consist of news and various radio shows in languages including French, Guerze, Koniaka, Kissi, Kono, Maninka, Mano, Pular, Susu, and Toma. Some radio shows include phone calls, background and foreground music, and various noise types. We collected this dataset for the purpose of unsupervized speech representation learning. A validation set of 300 tagged audio clips is also included. Please see our paper for more details on this dataset. Additional resources can be found in the following git repository: github.com

Visit

github.com

Connected records

paper

Tasks

automatic speech recognitiontext to speechspeech processing

Languages

Kisi, SouthernKonoManinka, KonyankaManinkakan, EasternManinkakan, WesternPularSusuTomaXaasongaxango

Tags

speech representation learning

Licenses

Creative Commons Attribution-ShareAlike 4.0 International License

Similaires

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech RecognitionThe Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech RecognitionSAE Radio News Speech CorpusAfrikaans Radio News Speech CorpusWest African Virtual Assistant Speech Recognition CorpusAn "African" Gospel: American Evangelical Radio in West Africa, 1954-1970

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

The Makerere AI Lab has built an end-to-end CTC Luganda ASR model using radio data. Having encountered data challenges in working with low resource languages, we take the initiative together with our partners to release the first radio corpus for Luganda. The corp

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

Building a usable radio monitoring automatic speech recognition (ASR) system is a challenging task for under-resourced languages and yet this is paramount in societies where radio is the main medium of public communication and discussions. Initial efforts by the Un

SAE Radio News Speech Corpus

News bulletins purchased from the SABC. Data to be used for the development of a large vocabulary co

Afrikaans Radio News Speech Corpus

News bulletins purchased from the SABC. Data to be used for the development of a large vocabulary co

West African Virtual Assistant Speech Recognition Corpus

This dataset contains 10,083 recorded utterances in French, Maninka, Pular and Susu from 49 speakers (16 female and 33 male) ranging from 5 to 76 years old on a variety of devices. Please see our paper for more details on this dataset. Additional resources can be

An "African" Gospel: American Evangelical Radio in West Africa, 1954-1970

During the second half of the twentieth century, Christianity underwent an epochal transformation fr