Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Bi-directional Recurrent End-to-End Neural Network Classifier for Spoken Arab Digit Recognition

Domain:

natural language processing

Record type:

papermodel
Creator:
ZerAbdBouRay
Editor:
UniCre
Publisher:
CCSDIEEE
Host:avatar
International audience —Automatic Speech Recognition can be considered as a transcription of spoken utterances into text which can be used to monitor/command a specific system. In this paper, we propose a general end-to-end approach to sequence learning that uses Long Short-Term Memory (LSTM) to deal with the non-uniform sequence length of the speech utterances. The neural architecture can recognize the Arabic spoken digit spelling of an isolated Arabic word using a classification methodology, with the aim to enable natural human-machine interaction. The proposed system consists to, first, extract the relevant features from the input speech signal using Mel Frequency Cepstral Coefficients (MFCC) and then these features are processed by a deep neural network able to deal with the non uniformity of the sequences length. A recurrent LSTM or GRU architecture is used to encode sequences of MFCC features as a fixed size vector that will feed a multilayer perceptron network to perform the classification. The whole neural network classifier is trained in an end-to-end manner. The proposed system outperforms by a large gap the previous published results on the same database.

Visit

inria.hal.science

Tasks

automatic speech recognitionspeech processing

Tags

Multilayer perceptron networkLong Short-Term MemorySpeech recognitionArabic digitsMel Frequency Cepstral CoefficientsAuto-encoder[INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL][INFO.INFO-HC]Computer Science [cs]/Human-Computer Interaction [cs.HC][INFO.INFO-NE]Computer Science [cs]/Neural and Evolutionary Computing [cs.NE]

Licenses

https://about.hal.science/hal-authorisation-v1/info:eu-repo/semantics/OpenAccess

Similar

Convolutional Neural Networks and Language Embeddings for End-to-End Dialect RecognitionEnd-to-End Mobile System for Diabetic Retinopathy Screening Based on Lightweight Deep Neural NetworkOkwuGbé: End-to-End Speech Recognition for Fon and IgboAn End-to-End Scene Text Recognition for Bilingual TextEnd-to-End Continuous Ethiopia Sign Language RecognitionEnd-To-End Continuous Ethiopian Sign Language Recognition

Convolutional Neural Networks and Language Embeddings for End-to-End Dialect Recognition

Dialect identification (DID) is a special case of general language identification (LID), but a more

End-to-End Mobile System for Diabetic Retinopathy Screening Based on Lightweight Deep Neural Network

End-to-End Mobile System for Diabetic Retinopathy Screening Based on Lightweight Deep Neural Network

Poster presented at the Deep Learning Indaba 2022 by Yaroub ELLOUMI

OkwuGbé: End-to-End Speech Recognition for Fon and Igbo

Language is inherent and compulsory for human communication. Whether expressed in a written or spoken way, it ensures understanding between people of the same and different regions. With the growing awareness and effort to include more low-resourced languages in NL

An End-to-End Scene Text Recognition for Bilingual Text

Text localization and recognition from natural scene images has gained a lot of attention recently d

End-to-End Continuous Ethiopia Sign Language Recognition

This paper tackles the challenge of continuous Ethiopian Sign Language (EthSL) recogn

End-To-End Continuous Ethiopian Sign Language Recognition