The Ethio Speech Corpus dataset is a multilingual speech corpus containing audio–text pairs across five Ethiopian languages.
It is designed to support the development of speech-to-text technologies for low-resource languages.
This dataset is part of the Afrivoices initiative — a collaborative effort to create a large-scale ASR dataset for African languages.