Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Enhancing Amharic Speech Recognition in Noisy Conditions through End-to-End Deep Learning

Domain:

natural language processing

Record type:

paper
Creator:
YohTes
Publisher:
MDP
Host:
Speech recognition, also known as automatic speech recognition (ASR), is a technology that enables software to transcribe spoken language into text. However, existing Amharic ASR methods require multiple separate blocks, such as language, acoustic, and pronunciation models with dictionaries, which can be time-consuming and influence performance. This study proposes an approach that replaces much of the speech pipeline with a single recurrent neural network (RNN) architecture. Our proposed architecture is based on a hybrid approach that combines a convolutional neural network (CNN) with a recurrent neural network (RNN) and a connectionist temporal classification (CTC) loss function. We conducted several experiments with noisy audio data that contain 20,000 valid sentences. The model was evaluated using the word error rate (WER) metric, achieving impressive results of 7% WER on noisy data. This approach has significant implications for the field of speech recognition, as it reduces the human effort required to create dictionaries and improves the efficiency and accuracy of ASR systems, making them more practical for real-world applications.

Visit

doi.org

Tasks

automatic speech recognitionspeech processing

Languages

Amharic

Licenses

http://creativecommons.org/licenses/by/4.0

Similar

End-to-End Amharic speech to text with Noisy data(Only sample data)A Noise-Robust End-to-End Framework for Amharic Speech RecognitionBreaking Barriers in Amharic Speech Recognition: A Scalable End-to-End ApproachEnd-to-End Historical Handwritten Ethiopic Text Recognition Using Deep LearningAn Investigation of End-to-End Multichannel Speech Recognition for Reverberant and Mismatch ConditionsSub-word Based End-to-End Speech Recognition for an Under-Resourced Language: Amharic

End-to-End Amharic speech to text with Noisy data(Only sample data)

This dataset is a nuanced compilation of diverse and noisy data meticulously curated to cater to the

A Noise-Robust End-to-End Framework for Amharic Speech Recognition

Abstract End-to-end automatic speech recognition (ASR) offers a streamlined altern

Breaking Barriers in Amharic Speech Recognition: A Scalable End-to-End Approach

End-to-End Historical Handwritten Ethiopic Text Recognition Using Deep Learning

An Investigation of End-to-End Multichannel Speech Recognition for Reverberant and Mismatch Conditions

Sequence-to-sequence (S2S) modeling is becoming a popular paradigm for automatic speech recognition

Sub-word Based End-to-End Speech Recognition for an Under-Resourced Language: Amharic

In this work, we focused on end-to-end speech recognition for less-resourced language, Amharic. The result can be integrated with other tasks such as spoken content retrieval. We explored three models, which consist of Convolutional Neural Networks, Recurrent Neura