Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Large Scale Speech Recognition for Low Resource Language Amharic, an End-to-End Approach

Domain:

natural language processing

Record type:

paper
Creator:
YohTes
Publisher:
MDP
Host:
Speech recognition, or automatic speech recognition (ASR), is a technology designed to convert spoken language into text using software. However, conventional ASR methods involve several distinct components, including language, acoustic, and pronunciation models with dictionaries. This modular approach can be time-consuming and may influence performance. In this study, we propose a method that streamlines the speech recognition process by incorporating a unified recurrent neural network (RNN) architecture. Our architecture integrates a convolutional neural network (CNN) with an RNN and employs a connectionist temporal classification (CTC) loss function. Key experiments were carried out using a dataset comprising 576,656 valid sentences, using erosion techniques. Evaluation of the model performance, measured by the word error rate (WER) metric, demonstrated remarkable results, achieving a WER of 2%. This approach has significant implications for the realm of speech recognition, as it alleviates the need for labor-intensive dictionary creation, enhancing the efficiency and accuracy of ASR systems, and making them more applicable to real-world scenarios. For future enhancements, we recommend the inclusion of dialectal and spontaneous data in the dataset to broaden the model's adaptability. Additionally, fine-tuning the model for specific tasks can optimize its performance for targeted objectives or domains, further enhancing its effectiveness in those areas.

Visit

doi.org

Tasks

automatic speech recognitionspeech processing

Languages

Amharic

Licenses

http://creativecommons.org/licenses/by/4.0

Similar

Sub-word Based End-to-End Speech Recognition for an Under-Resourced Language: AmharicEnd-to-End Speech Recognition with Deep Fusion: Leveraging External Language Models for Low-Resource ScenariosBreaking Barriers in Amharic Speech Recognition: A Scalable End-to-End ApproachA Noise-Robust End-to-End Framework for Amharic Speech RecognitionImproving End-to-End Speech Translation for the Low Resource Language Fongbe to FrenchDziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect

Sub-word Based End-to-End Speech Recognition for an Under-Resourced Language: Amharic

In this work, we focused on end-to-end speech recognition for less-resourced language, Amharic. The result can be integrated with other tasks such as spoken content retrieval. We explored three models, which consist of Convolutional Neural Networks, Recurrent Neura

End-to-End Speech Recognition with Deep Fusion: Leveraging External Language Models for Low-Resource Scenarios

With the rapid development of Automatic Speech Recognition (ASR) technology, end-to-end speech recog

Breaking Barriers in Amharic Speech Recognition: A Scalable End-to-End Approach

A Noise-Robust End-to-End Framework for Amharic Speech Recognition

Abstract End-to-end automatic speech recognition (ASR) offers a streamlined altern

Improving End-to-End Speech Translation for the Low Resource Language Fongbe to French

This study addresses the challenges of end-to-end (E2E) Speech-to-Text Translation (STT) for the low

Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect

Automatic speech and language technologies are still heavily biased toward high-resource languages,