Afrivoice is a robust audio processing pipeline designed to transcribe speech from African languages, translate it (via English), and synthesize the translated text back into audio. It integrates a wide range of ASR and TTS models to support African linguistic diversity.
# multilingual-african-speech-translation
Afrivoice is a robust audio processing pipeline designed to transcribe speech from African languages, translate it (via English), and synthesize the translated text back into audio. It integrates a wide range of ASR and TTS models to support African linguistic diversity.
Features
🎙️ Automatic Speech Recognition (ASR) for over 20 African languages using models from SpeechBrain, HuggingFace, and Whisper.
🌐 Translation via Microsoft Azure or Google Translate APIs (configurable).
🔊 Text-to-Speech (TTS) for supported languages using VITS models.
🔄 End-to-End Pipeline: Input speech in one language → Transcription → Translation → Synthesized audio in another language.
Supported Languages
Language Code Language ASR Support TTS Support
kin Kinyarwanda ✅ ✅
lug Luganda ✅ ✅
swh Swahili ✅ ✅
am Amharic ✅ ✅
ha Hausa ✅ 🚫*
yo Yoruba ✅ ✅
ig Igbo ✅ 🚫*
xho Xhosa ✅ ✅
zul Zulu ✅ ✅
tir Tigrinya ✅ ✅
so Somali ✅ ✅
ach, nyn, teo, lgg Sunbird languages (Acholi, Runyankole, Ateso, Lugbara) ✅ ✅