Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Edge-Based Speech Transcription and Synthesis for Kinyarwanda and Swahili Languages

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
MboTuyBiyTon
Hôte:avatar
This paper presents a novel framework for speech transcription and synthesis, leveraging edge-cloud parallelism to enhance processing speed and accessibility for Kinyarwanda and Swahili speakers. It addresses the scarcity of powerful language processing tools for these widely spoken languages in East African countries with limited technological infrastructure. The framework utilizes the Whisper and SpeechT5 pre-trained models to enable speech-to-text (STT) and text-to-speech (TTS) translation. The architecture uses a cascading mechanism that distributes the model inference workload between the edge device and the cloud, thereby reducing latency and resource usage, benefiting both ends. On the edge device, our approach achieves a memory usage compression of 9.5% for the SpeechT5 model and 14% for the Whisper model, with a maximum memory usage of 149 MB. Experimental results indicate that on a 1.7 GHz CPU edge device with a 1 MB/s network bandwidth, the system can process a 270-character text in less than a minute for both speech-to-text and text-to-speech transcription. Using real-world survey data from Kenya, it is shown that the cascaded edge-cloud architecture proposed could easily serve as an excellent platform for STT and TTS transcription with good accuracy and response time.

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processingtext to speech

Languages

KinyarwandaSwahili

Tags

Distributed, Parallel, and Cluster ComputingMachine Learning

Similaires

Multilingual Automatic Speech Recognition for Kinyarwanda, Swahili, and Luganda: Advancing ASR in Select East African LanguagesA Curated Crowdsourced Dataset of Luganda and Swahili Speech for Text-to-Speech Synthesissamolubukun/Nigerian-Languages-WAZOBIA-Speech-TranscriptionText to speech synthesis for ethiopian semitic languages: Issues and the way forwardFast transcription of speech in low-resource languagesStatistical modelling of speech units in HMM-based speech synthesis for Arabic

Multilingual Automatic Speech Recognition for Kinyarwanda, Swahili, and Luganda: Advancing ASR in Select East African Languages

Multilingual Automatic Speech Recognition for Kinyarwanda, Swahili, and Luganda: Advancing ASR in Select East African Languages

Poster presented at the Deep Learning Indaba 2023 by Samuel Rutunda

A Curated Crowdsourced Dataset of Luganda and Swahili Speech for Text-to-Speech Synthesis

This dataset contains curated and preprocessed speech recordings in Luganda and Kiswahili for use in

samolubukun/Nigerian-Languages-WAZOBIA-Speech-Transcription

This Gradio application provides a user-friendly interface for transcribing spoken audio in three ma

Text to speech synthesis for ethiopian semitic languages: Issues and the way forward

Fast transcription of speech in low-resource languages

We present software that, in only a few hours, transcribes forty hours of recorded speech in a surpr

Statistical modelling of speech units in HMM-based speech synthesis for Arabic

International audience This paper investigates statistical parametric speech synthesi