Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

hresholding in Convolutional Neural Network Model for Yorùbá Speech-to-Text Conversion

Domaine:

natural language processing

Type de record:

modelpaper
Créateur:
SunJohn. B. OladosuAdeAde
Éditeur:
Afr
Hôte:
Speech-to-text and text-to-speech conversions are referred to as Automatic Speech Recognition (ASR), which employs transcripts to transform speech signals into text.  Convolutional Neural Networks (CNNs) have considerable efficacy in ASR, particularly in the feature extraction stage, where they outperform traditional models in capturing local hierarchical features from spectrogram representations of audio signals.  Nonetheless, the efficacy of standard CNN across multiple thresholds in speech-to-text conversion for low-resource languages remains under investigation.  This study examines the effects of two distinct thresholds (0.22 and 0.35) in a CNN-based model for Yorùbá speech-to-text conversion.  A corpus of 64 unique Yorùbá language samples was utilized for implementation and evaluation in MATLAB R2023a.   The model is evaluated using False Positive Rate (FPR), Specificity (Spec), Sensitivity (Sen), Precision (Prec), Accuracy (Acc), and F1-score.  The FPR diminishes by 0.19%, but the SPEC ascends by 0.19%, signifying enhanced true negative detection at the elevated threshold. The speech-to-text conversion of the Yorùbá language dataset was precise, with a reduction in SEN of 0.12%.  The PREC climbs by 0.22% as the threshold increases, indicating an enhancement in the model's accuracy for affirmative case predictions.  An increase of 0.5% in the Acc signifies an enhancement in overall performance.   A more balanced performance is evidenced by a 0.05% enhancement in the F1 Score, which equilibrates precision and recall. Performance metrics reveal a modest enhancement of 0.35. The results underscore the capacity of language technology to embody global linguistic diversity and enhance the accessibility of speech-to-text conversion systems.

Visit

doi.org

Tasks

automatic speech recognitionspeech processing

Languages

Yoruba

Licenses

https://creativecommons.org/licenses/by-nc/4.0

Similaires

Yorùbá Character Recognition System Using Convolutional Recurrent Neural NetworkA Convolutional Neural Network Model for Setswana Named Entity RecognitionAn Innovative Word Encoding Method For Text Classification Using Convolutional Neural NetworkResearch on Deep Neural Network for Afaan-Oromo Language Text-to-Speech SynthesisGRAPH CONVOLUTIONAL NETWORK AND RECURRENT NEURAL NETWORK ENSEMBLE FOR EXTRACTIVE TEXT SUMMARISATION IN THE HAUSA LANGUAGEConvolutional neural network for speech emotion recognition in the Moroccan Arabic dialect language

Yorùbá Character Recognition System Using Convolutional Recurrent Neural Network

Handwritten recognition systems enable automatic recognition of human handwritings, thereby increasi

A Convolutional Neural Network Model for Setswana Named Entity Recognition

Named entity recognition (NER) is a key component of the core task of natural language processing (N

An Innovative Word Encoding Method For Text Classification Using Convolutional Neural Network

Text classification plays a vital role today especially with the intensive use of social networking

Research on Deep Neural Network for Afaan-Oromo Language Text-to-Speech Synthesis

GRAPH CONVOLUTIONAL NETWORK AND RECURRENT NEURAL NETWORK ENSEMBLE FOR EXTRACTIVE TEXT SUMMARISATION IN THE HAUSA LANGUAGE

Automatic Text Summarisation (ATS) is crucial for managing information overload, especially in low-r

Convolutional neural network for speech emotion recognition in the Moroccan Arabic dialect language

Extracting the speaker's emotional state has become an active research topic lately due to the deman