Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

West African Radio Corpus

Domain:

natural language processing

Record type:

dataset
This dataset contains 17,090 audio clips of length 30 seconds sampled from archives collected from 6 Guinean radio stations. The broadcasts consist of news and various radio shows in languages including French, Guerze, Koniaka, Kissi, Kono, Maninka, Mano, Pular, Susu, and Toma. Some radio shows include phone calls, background and foreground music, and various noise types. We collected this dataset for the purpose of unsupervized speech representation learning. A validation set of 300 tagged audio clips is also included. Please see our paper for more details on this dataset. Additional resources can be found in the following git repository: github.com

Visit

github.com

Connected records

paper

Tasks

automatic speech recognitiontext to speechspeech processing

Languages

Kisi, SouthernKonoManinka, KonyankaManinkakan, EasternManinkakan, WesternPularSusuTomaXaasongaxango

Tags

speech representation learning

Licenses

Creative Commons Attribution-ShareAlike 4.0 International License