This dataset contains audio clips of Kalenjin speech and their corresponding transcriptions. It is intended for use in training and evaluating Automatic Speech Recognition (ASR) models for the Kalenjin language.
This dataset was created from the Mozilla Common Voice project. It contains a total of [NUMBER] hours of audio, split into train, test, and validated sets.
Languages