Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

hundehanna/amharic-asr

Domaine:

natural language processing

Type de record:

model
Créateur:
hun
Hôte:
Fine-tuned OpenAI Whisper-medium on the Leyu Amharic dataset (4 dialects, ~27k samples) for Automatic Speech Recognition. Includes Colab training notebook and Gradio demo. # Amharic Speech Recognition — Fine-tuned Whisper Medium Fine-tuned `openai/whisper-medium` on the Leyu Amharic dataset — a multi-dialect Amharic speech corpus covering Gojjam, Gonder, Wello, and Shewa dialects (~27,000 samples). This project demonstrates end-to-end low-resource ASR development for Amharic (አማርኛ), one of Ethiopia's most widely spoken languages. --- ## Live Demo > Try the model directly in your browser — no setup required. **🎙️ Launch Demo on HuggingFace Spaces** Record your voice or upload an audio file to receive an Amharic transcript. --- ## Results | Model | WER ↓ | Notes | |-------|--------|-------| | `openai/whisper-medium` (zero-shot baseline) | _TBD_ | No fine-tuning | | `whisper-medium-amharic` (this model) | _TBD_ | Fine-tuned on Leyu | > Results will be updated after training completes. --- ## Dataset: Leyu Amharic | Dialect | HuggingFace ID | Samples | Duration | |---------|----------------|---------|----------| | Gojjam | `leyu-amharic/leyu-amharic-gojjam-dialect` | 10,575 | 81.6h | | Gonder | `leyu-amharic/leyu-amharic-gonder-dialect` | 8,990 | — | | Wello | `leyu-amharic/leyu-amharic-wello-dialect` | 4,860 | — | | Shewa | `leyu-amharic/leyu-amharic-shewa-dialect` | 2,590 | — | | **Total** | | **~27,000** | | - Audio: `.wav`, 16kHz mono - Transcripts: Ethiopic script (`text` column) - Speakers: mixed gender, recorded on mobile devices in real environments - All datasets only ship with a `train` split — we apply an 80/10/10 train/val/test split --- ## Model - **Base**: `openai/whisper-medium` (307M parameters) - **Task**: Automatic Speech Recognition (`transcribe`) - **Language**: Amharic (`am`) - **Training**: Google Colab T4 GPU, fp16, gradient checkpointing - **Framework**: HuggingFace `transformers` + `Seq2SeqTrainer` --- ## Project Structure ``` amharic-asr/ ├── notebooks/ │ └── train_whisper_amharic.ipynb # Colab training notebook (runnable) ├── src/ │ ├── data_prep.py # Dat …

Visit

github.com

Languages

AmharicGeez

Licenses

MIT

Similaires

mintesnot96/amharic-asribso99/Amharic-ASR-agkphysics/amharic-asryamlakyam/Amharic-ASRbeimnet777/amharic-asrIsraelAbebe/Amharic-ASR-Dataset

mintesnot96/amharic-asr

amharic-asr, amharic STT # Amharic Speech Recognition — Fine-tuned Whisper Medium Fine-tuned `

ibso99/Amharic-ASR-

Am Amharic speech to text or ASR model trained on 205.17 hours of data # Amharic-ASR- # Amharic AS

agkphysics/amharic-asr

Amharic wav2vec 2,0 ASR model trained with HuggingFace Transformers # Amharic ASR Amharic wav2vec 2

yamlakyam/Amharic-ASR

**Introduction:** This repository contains code for training an Automatic Speech Recognition (ASR) s

beimnet777/amharic-asr

ALFFA_PUBLIC #####If you use this data, please cite the following paper for the ressources @article{

IsraelAbebe/Amharic-ASR-Dataset

speech recognition for Amharic language # Amharic-ASR-Dataset speech recognition for Amharic langua