This dataset contains transcribed Kinyarwanda audio, designed to support training and evaluation of Automatic Speech Recognition (ASR) systems. It is part of a study on how varying training data volumes affect model performance using Whisper-large-v3.
The full dataset consists of approximately 263,000 audio samples covering 5 key domains:
🏥 Health
🏛️ Government
💰 Financial Services
🎓 Education
🌾 Agriculture