Logo Lanfrica

Waxal Speech Data Resources

Domain:

natural language processing

Record type:

dataset

The Waxal Speech Data Resources project aims to create natural language processing (NLP) resources for diverse African languages using crowdsourced speech and text data. The goal is to develop robust multilingual NLP systems (Speech, Neural Machine Translation, Question & Answer, Language Models) that can handle variations in accents and code-switching, and are optimized for low-end mobile devices. This initiative seeks to make NLP more inclusive and advance deep learning for NLP and machine learning under memory/computing constraints.

LanguageParticipantsRecordingsSpeech hoursTranscribed Hours
Wolof4242862965196.45

Currently, the project includes Wolof data from 4,242 participants, totaling 86,296 recordings, 519 speech hours, and 6.45 transcribed hours.

The project addresses the significant disparity in NLP system performance for African languages compared to European languages, despite a similar or larger number of speakers. It also acknowledges the substantial differences between standard and local variations of official European languages spoken in Africa.

The data, including audio files, response tables, transcription tables, and translation tables, can be downloaded from the project's GitHub releases. Audio data was collected via a chatbot using WhatsApp and Twilio, while transcription and translation were done by linguists. Instructions for setting up independent speech data collection are also provided.