This dataset contains processed Kinyarwanda speech data with trimmed audio segments.
The dataset contains two splits:
dev_test: 9,263 samples
test: 9,265 samples
Each sample contains:
id: Unique identifier for the sample
audio: Audio data
audio_language: Language of the audio (Kinyarwanda)
text: Transcription of the audio
prompt: Associated prompt or context