This repository contains curated subsets of the Digital Umuganda / AfriVoices dataset for the Shona, Lingala, Fulani, and Malagasy languages. The dataset is split into train and test sets for each language.
Train split: Contains audio clips with corresponding transcriptions.
Test split: Contains audio clips without transcriptions, as none were available in the original source.
Each sample in the dataset includes the following fields: