Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

West African Virtual Assistant Speech Recognition Corpus

Domain:

natural language processing

Record type:

dataset
This dataset contains 10,083 recorded utterances in French, Maninka, Pular and Susu from 49 speakers (16 female and 33 male) ranging from 5 to 76 years old on a variety of devices. Please see our paper for more details on this dataset. Additional resources can be found in the following git repository: github.com

Visit

github.com

Tasks

automatic speech recognitionspeech processing

Languages

Maninkakan, EasternManinkakan, WesternPularSusuXaasongaxango

Licenses

Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)

Similar

Using Radio Archives for Low-Resource Speech Recognition: Towards an Intelligent Virtual Assistant for Illiterate UsersThe Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech RecognitionThe Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognitionmumbi12337/icu-virtual-assistantMDVC corpus: empowering Moroccan Darija speech recognitionDesign of a Tigrinya Language Speech Corpus for Speech Recognition

Using Radio Archives for Low-Resource Speech Recognition: Towards an Intelligent Virtual Assistant for Illiterate Users

For many of the 700 million illiterate people around the world, speech recognition technology could provide a bridge to valuable information and services. Yet, those most in need of this technology are often the most underserved by it. In many countries, illiterate

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

The Makerere AI Lab has built an end-to-end CTC Luganda ASR model using radio data. Having encountered data challenges in working with low resource languages, we take the initiative together with our partners to release the first radio corpus for Luganda. The corp

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

Building a usable radio monitoring automatic speech recognition (ASR) system is a challenging task for under-resourced languages and yet this is paramount in societies where radio is the main medium of public communication and discussions. Initial efforts by the Un

mumbi12337/icu-virtual-assistant

This is a smart, interactive chatbot widget built completely from scratch for the Information and Co

MDVC corpus: empowering Moroccan Darija speech recognition

Automatic speech recognition (ASR) technology has significantly transformed human-machine interactio

Design of a Tigrinya Language Speech Corpus for Speech Recognition

In this paper, we describe the first Tigrinya Languages speech corpora designed and development for speech recognition purposes. Tigrinya, often written as Tigrigna (ትግርኛ) /tɪˈɡrinjə/ belongs to the Semitic branch of the Afro-Asiatic languages where it shows the ch