Can you create an automatic speech recognition (ASR) model for African accents, for use by doctors?
African hospitals have some of the lowest doctor-patient ratios in the world. At very busy clinics, doctors could see over 30 patients a day without any of the productivity-boosting tools available to their colleagues in developed countries.
Clinical speech-to-text is ubiquitous in the developed world but virtually absent across African hospitals. This project seeks to create pan-African English ASR models for healthcare, expanding clinical speech recognition access to African clinics to help alleviate the burden of daily clinical documentation.
The objective of this challenge is to create an automatic speech recognition model for clinical speech-to-text, using an accented English speech corpus of 200 hours, featuring 120 different African accents.
This is the largest and most diverse open-source accented speech dataset for clinical and general domain ASR in Africa, covering 13 countries, 2463 unique speakers, 52% female; featuring 67,577 audio clips, and 200 hours of audio.
About the Data
The data is from two domains, healthcare (~60%) and general (~40%); general domain includes news, sports, entertainment, politics, and Wikipedia.
There are 196 hours of accented English recordings; audio clips are ~ 11 seconds on average and are from 13 different countries covering 120 accents from West, South, and East Africa.
There are 57 819 recordings in train, 3 227 in dev and 5 070 test.