This dataset contains Sheng audio and accurate human-transcript pairs within the mobile money domain. It covers a wide range of demographics with a rich distribution across gender/age group pairs. Audio was collected from individual respondents in-field (with phone microphone as the most common recording methodology).
Speakers were asked to answer prompts/questions that focused specifically on mobile money usage, and provided free-form, unstructured responses. Rather than reading predetermined sentences, respondents spoke naturally about their experiences and usage. This methodology captures a diverse range of speakers, utterances, speaking styles, and recording environments.
This makes the dataset ideal for the development of ASR modelling and applications where code-switching between Sheng, Swahili and English is expected, and speech is likely to occur in real-world environments where there is high acoustic variability.
Currently only the 1 hour sample is available for download; we have a much larger corpus of data that will be published as a final dataset of 50 hours.