Each datapoint in this dataset consists of a JPEG image, a corresponding audio Webm file describing
This is an image prompt ASR dataset for Swahili. The dataset was collected on 5 domains: Agriculture
Language Category Total number of hours Total number of transcribed hours Total number of clips Tota
Domain Total number of hours Total number of transcribed hours Total number of clips Total Size of t