Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Amharic ASR Dataset (CC-BY 4.0)

Domain:

natural language processing

Record type:

dataset
Creator:
cha
Host:
This dataset is a processed version of the original BDU-speech dataset by Yohannes A. Ejigu. It contains paired Amharic speech audio and transcriptions, structured for use in automatic speech recognition (ASR) research and model training. Audio files are decoded as mono. Sampling rates may vary across files. DatasetDict({ train: Dataset({ features: ['audio', 'sentence'],

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Languages

Amharic

Licenses

cc-by-4.0