Logo Lanfrica

Amharic Speech Corpus

Domain:

natural language processing

Record type:

dataset
This is an Amharic speech corpus which is suitable for the development and evaluation of speech recognition and retrieval systems. The corpus contains 110 hours of speech data with syllable and grapheme-based transcriptions collected from public domain or resources with specific permissive licenses. The corpus is partitioned into training and validation set which contains smaller audio segments not longer than 28 seconds. Utterances in each partition are re-sampled with a sampling frequency of 16 kHz with a sample size of 16 bits, 256kbs bitrate with a mono channel and stored as a wav file. The syllable and grapheme-based transcriptions are provided in plain text and including the audio details as a json and csv files. For more details about the corpus, refer to the original publication.