KasaSpeech is a large-scale English–Twi code-switching speech dataset created to advance speech AI
research for low-resource African languages. The dataset contains transcribed speech recordings
featuring natural switching between English and Twi, collected from diverse speakers primarily in Ghana.
The current version of the dataset shows these statistics
Split
Samples
Train