Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

End-to-End Speech Recognition with Deep Fusion: Leveraging External Language Models for Low-Resource Scenarios

Domaine:

natural language processing

Type de record:

paper
Créateur:
LusShiZho
Éditeur:
MDP
Hôte:
With the rapid development of Automatic Speech Recognition (ASR) technology, end-to-end speech recognition systems have gained significant attention due to their ability to directly convert raw speech signals into text. However, such systems heavily rely on large amounts of labeled speech data, which severely limits model training performance and generalization, especially in low-resource language environments. To address this issue, this paper proposes an end-to-end speech recognition approach based on deep fusion, which tightly integrates an external language model (LM) with the end-to-end model during the training phase, effectively compensating for the lack of linguistic prior knowledge. Unlike traditional shallow fusion methods, deep fusion enables the model and the external LM to share representations and jointly optimize during training, thereby enhancing recognition performance under low-resource conditions. Experiments conducted on the Common Voice dataset show that, in a 10 h extremely low-resource scenario, the deep fusion method reduces the character error rate (CER) from 51.1% to 17.65%. In a 100 h scenario, it achieves a relative reduction of approximately 2.8%. Furthermore, ablation studies on model layers demonstrate that even with a reduced number of encoder and decoder layers to decrease model complexity, deep fusion continues to effectively leverage external linguistic priors, significantly improving performance in low-resource speech recognition tasks.

Visit

doi.org

Tasks

automatic speech recognitionspeech processing

Licenses

https://creativecommons.org/licenses/by/4.0/

Similaires

Large Scale Speech Recognition for Low Resource Language Amharic, an End-to-End ApproachImproving End-to-End Speech Translation for the Low Resource Language Fongbe to FrenchMultilingual Speech Recognition With A Single End-To-End ModelDziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian DialectEnhancing Amharic Speech Recognition in Noisy Conditions through End-to-End Deep LearningOkwuGbé: End-to-End Speech Recognition for Fon and Igbo

Large Scale Speech Recognition for Low Resource Language Amharic, an End-to-End Approach

Speech recognition, or automatic speech recognition (ASR), is a technology designed to convert spoke

Improving End-to-End Speech Translation for the Low Resource Language Fongbe to French

This study addresses the challenges of end-to-end (E2E) Speech-to-Text Translation (STT) for the low

Multilingual Speech Recognition With A Single End-To-End Model

Training a conventional automatic speech recognition (ASR) system to support multiple languages is c

Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect

Automatic speech and language technologies are still heavily biased toward high-resource languages,

Enhancing Amharic Speech Recognition in Noisy Conditions through End-to-End Deep Learning

Speech recognition, also known as automatic speech recognition (ASR), is a technology that enables s

OkwuGbé: End-to-End Speech Recognition for Fon and Igbo

Language is inherent and compulsory for human communication. Whether expressed in a written or spoken way, it ensures understanding between people of the same and different regions. With the growing awareness and effort to include more low-resourced languages in NL