Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Advanced Convolutional Neural Network-Based Hybrid Acoustic Models for Low-Resource Speech Recognition

Domain:

natural language processing

Record type:

paper
Creator:
TesJunTul
Publisher:
MDP
Host:
Deep neural networks (DNNs) have shown a great achievement in acoustic modeling for speech recognition task. Of these networks, convolutional neural network (CNN) is an effective network for representing the local properties of the speech formants. However, CNN is not suitable for modeling the long-term context dependencies between speech signal frames. Recently, the recurrent neural networks (RNNs) have shown great abilities for modeling long-term context dependencies. However, the performance of RNNs is not good for low-resource speech recognition tasks, and is even worse than the conventional feed-forward neural networks. Moreover, these networks often overfit severely on the training corpus in the low-resource speech recognition tasks. This paper presents the results of our contributions to combine CNN and conventional RNN with gate, highway, and residual networks to reduce the above problems. The optimal neural network structures and training strategies for the proposed neural network models are explored. Experiments were conducted on the Amharic and Chaha datasets, as well as on the limited language packages (10-h) of the benchmark datasets released under the Intelligence Advanced Research Projects Activity (IARPA) Babel Program. The proposed neural network models achieve 0.1–42.79% relative performance improvements over their corresponding feed-forward DNN, CNN, bidirectional RNN (BRNN), or bidirectional gated recurrent unit (BGRU) baselines across six language collections. These approaches are promising candidates for developing better performance acoustic models for low-resource speech recognition tasks.

Visit

doi.org

Tasks

automatic speech recognitionspeech processing

Languages

AmharicSebat Bet Gurage

Licenses

https://creativecommons.org/licenses/by/4.0/

Similar

Syllable-Based and Hybrid Acoustic Models for Amharic Speech RecognitionSpeech recognition system based on deep neural network acoustic modeling for low resourced language-AmharicAdaptive Activation Network For Low Resource Multilingual Speech RecognitionImproving Arabic handwritten text recognition through transfer learning with convolutional neural network-based modelsNeural Network Based Hausa Language Speech RecognitionConvolutional neural network for speech emotion recognition in the Moroccan Arabic dialect language

Syllable-Based and Hybrid Acoustic Models for Amharic Speech Recognition

International audience no abstract

Speech recognition system based on deep neural network acoustic modeling for low resourced language-Amharic

Adaptive Activation Network For Low Resource Multilingual Speech Recognition

Low resource automatic speech recognition (ASR) is a useful but thorny task, since deep learning ASR

Improving Arabic handwritten text recognition through transfer learning with convolutional neural network-based models

Arabic handwritten text recognition is a complex and challenging research domain. This study propose

Neural Network Based Hausa Language Speech Recognition

Convolutional neural network for speech emotion recognition in the Moroccan Arabic dialect language

Extracting the speaker's emotional state has become an active research topic lately due to the deman