Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Afrivoice ASR Swahili dataset

Domain:

natural language processing

Record type:

dataset
Creator:
Dig
Host:
This is an image prompt ASR dataset for Swahili. The dataset was collected on 5 domains: Agriculture, Education, Finance, Government and Health. Domain Total Hours Transcribed Hours Number of Clips Dataset Size (GB) Agriculture 739.67 717.66 128578 43.537 Education 414.3 398.1 71823 30.5219 Financial 615.29 604.78 106387 57.8098 Government 591.1 579.11 102915 49.6434

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Languages

Swahili

Tags

DigitalUmugandaDUswswaSwahiliasrsttvoicespeech

Licenses

cc-by-4.0

Similar

Afrivoice ASR Swahili datasetAfrivoice Kinyarwanda ASR datasetAfrivoice ASR kinyarwanda datasetAfrivoice ASR Ethiopia datasetDigitalUmuganda/Mbaza-ASR-Afrivoice-swahili-400hoursLazanantenaina/swahili-asr-dataset

Afrivoice ASR Swahili dataset

Domain Total number of hours Total number of transcribed hours Total number of clips Total Size of t

Afrivoice Kinyarwanda ASR dataset

Each datapoint in this dataset consists of a JPEG image, a corresponding audio Webm file describing

Afrivoice ASR kinyarwanda dataset

[need more information] [need more information] [need more information] [need more information]

Afrivoice ASR Ethiopia dataset

Language Category Total number of hours Total number of transcribed hours Total number of clips Tota

DigitalUmuganda/Mbaza-ASR-Afrivoice-swahili-400hours

Lazanantenaina/swahili-asr-dataset