Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Edge-Based Speech Transcription and Synthesis for Kinyarwanda and Swahili Languages

Domain:

natural language processing

Record type:

papermodel
Creator:
MboTuyBiyTon
Host:avatar
This paper presents a novel framework for speech transcription and synthesis, leveraging edge-cloud parallelism to enhance processing speed and accessibility for Kinyarwanda and Swahili speakers. It addresses the scarcity of powerful language processing tools for these widely spoken languages in East African countries with limited technological infrastructure. The framework utilizes the Whisper and SpeechT5 pre-trained models to enable speech-to-text (STT) and text-to-speech (TTS) translation. The architecture uses a cascading mechanism that distributes the model inference workload between the edge device and the cloud, thereby reducing latency and resource usage, benefiting both ends. On the edge device, our approach achieves a memory usage compression of 9.5% for the SpeechT5 model and 14% for the Whisper model, with a maximum memory usage of 149 MB. Experimental results indicate that on a 1.7 GHz CPU edge device with a 1 MB/s network bandwidth, the system can process a 270-character text in less than a minute for both speech-to-text and text-to-speech transcription. Using real-world survey data from Kenya, it is shown that the cascaded edge-cloud architecture proposed could easily serve as an excellent platform for STT and TTS transcription with good accuracy and response time.

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processingtext to speech

Languages

KinyarwandaSwahili

Tags

Distributed, Parallel, and Cluster ComputingMachine Learning

Similar

Multilingual Automatic Speech Recognition for Kinyarwanda, Swahili, and Luganda: Advancing ASR in Select East African LanguagesA Curated Crowdsourced Dataset of Luganda and Swahili Speech for Text-to-Speech Synthesissamolubukun/Nigerian-Languages-WAZOBIA-Speech-TranscriptionText to speech synthesis for ethiopian semitic languages: Issues and the way forwardFast transcription of speech in low-resource languagesStatistical modelling of speech units in HMM-based speech synthesis for Arabic

Multilingual Automatic Speech Recognition for Kinyarwanda, Swahili, and Luganda: Advancing ASR in Select East African Languages

Multilingual Automatic Speech Recognition for Kinyarwanda, Swahili, and Luganda: Advancing ASR in Select East African Languages

Poster presented at the Deep Learning Indaba 2023 by Samuel Rutunda

A Curated Crowdsourced Dataset of Luganda and Swahili Speech for Text-to-Speech Synthesis

This dataset contains curated and preprocessed speech recordings in Luganda and Kiswahili for use in

samolubukun/Nigerian-Languages-WAZOBIA-Speech-Transcription

This Gradio application provides a user-friendly interface for transcribing spoken audio in three ma

Text to speech synthesis for ethiopian semitic languages: Issues and the way forward

Fast transcription of speech in low-resource languages

We present software that, in only a few hours, transcribes forty hours of recorded speech in a surpr

Statistical modelling of speech units in HMM-based speech synthesis for Arabic

International audience This paper investigates statistical parametric speech synthesi