Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

evie-8/afrivoices

Domain:

natural language processing

Record type:

dataset
Creator:
evi
Host:
This repository contains curated subsets of the Digital Umuganda / AfriVoices dataset for the Shona, Lingala, Fulani, and Malagasy languages. The dataset is split into train and test sets for each language. Train split: Contains audio clips with corresponding transcriptions. Test split: Contains audio clips without transcriptions, as none were available in the original source. Each sample in the dataset includes the following fields:

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Languages

FulaFulfulde, AdamawaFulfulde, BorguFulfulde, Central-Eastern NigerFulfulde, MaasinaFulfulde, NigerianFulfulde, Western NigerLingalaMalagasyMalagasy, Merina+2

Similar

evie-8/afrivoices-imagesevie-8/kikuyu-dataevie-8/zulu-dataevie-8/kinyarwanda-speech-hackathonevie-8/swahili-speech-data

evie-8/afrivoices-images

This repository contains image-only subsets from the original Digital Umuganda / AfriVoices project.

evie-8/kikuyu-data

This dataset contains Kikuyu speech audio with transcriptions, sourced and adapted from the ANV Data

evie-8/zulu-data

The transcriptions in this dataset were generated using the TheirStory/whisper-medium-zulu model. Th

evie-8/kinyarwanda-speech-hackathon

This dataset contains transcribed Kinyarwanda audio, designed to support training and evaluation of

evie-8/swahili-speech-data

This dataset contains Swahili speech audio paired with transcriptions.It is split into training, dev