Somali ASR Corpus is a cleaned Somali speech dataset designed for Automatic Speech Recognition (ASR) and Speech-to-Text (STT) research. The dataset contains transcribed Somali speech recordings processed with noise reduction and audio normalization techniques to improve training quality for speech models.
Feature
Type
text
string
audio
Audio
The audio files were processed using: