SautiDB-Naija: A Nigerian L2 English Speech Corpus Poster presented at the Deep Learning Indaba 2022 by Tejumande Afonja
This paper describes foundational efforts with SautiDB-Naija, a novel corpus of non-native (L2) Nigerian English speech. We describe how the corpus was created and curated as well as preliminary experiments with accent classification and learning Nigerian accent em
The SautiDB dataset collection project is an ongoing effort to collect datasets of various Nigerian accents. The dataset was collected in an uncontrolled manner, users who visit our webapp can record their voice and contribute to the dataset. The webapp uses the au
African Speech Technology speech and transcription data for the English-English database. The "spee
Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test sui
African Speech Technology speech and transcription data for the Afrikaans-English database. The "sp
African Speech Technology speech and transcription data for the Coloured English database. The "spe