Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Identifying the influence of transfer learning method in developing an end-to-end automatic speech recognition system with a low data level

Domain:

natural language processing

Record type:

paper
Creator:
OrkKeyDinAkb
Publisher:
Pri
Host:
Ensuring the best quality and performance of modern speech technologies, today, is possible based on the widespread use of machine learning methods. The idea of this project is to study and implement an end-to-end system of automatic speech recognition using machine learning methods, as well as to develop new mathematical models and algorithms for solving the problem of automatic speech recognition for agglutinative (Turkic) languages. Many research papers have shown that deep learning methods make it easier to train automatic speech recognition systems that use an end-to-end approach. This method can also train an automatic speech recognition system directly, that is, without manual work with raw signals. Despite the good recognition quality, this model has some drawbacks. These disadvantages are based on the need for a large amount of data for training. This is a serious problem for low-data languages, especially Turkic languages such as Kazakh and Azerbaijani. To solve this problem, various methods are needed to apply. Some methods are used for end-to-end speech recognition of languages belonging to the group of languages of the same family (agglutinative languages). Method for low-resource languages is transfer learning, and for large resources – multi-task learning. To increase efficiency and quickly solve the problem associated with a limited resource, transfer learning was used for the end-to-end model. The transfer learning method helped to fit a model trained on the Kazakh dataset to the Azerbaijani dataset. Thereby, two language corpora were trained simultaneously. Conducted experiments with two corpora show that transfer learning can reduce the symbol error rate, phoneme error rate (PER), by 14.23 % compared to baseline models (DNN+HMM, WaveNet, and CNC+LM). Therefore, the realized model with the transfer method can be used to recognize other low-resource languages.

Visit

doi.org

Tasks

automatic speech recognitionspeech processingtransfer learning

Licenses

http://creativecommons.org/licenses/by/4.0

Similar

Towards End-to-End Training of Automatic Speech Recognition for Nigerian PidginMultilingual Speech Recognition With A Single End-To-End ModelEnd-to-End Automatic Speech Translation of AudiobooksLarge Scale Speech Recognition for Low Resource Language Amharic, an End-to-End ApproachDziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian DialectEnd-To-End Multilingual Automatic Speech Recognition For Less-Resourced Languages: The Case Of Four Ethiopian Languages

Towards End-to-End Training of Automatic Speech Recognition for Nigerian Pidgin

Nigerian Pidgin remains one of the most popular languages in West Africa. With at least 75 million speakers along the West African coast, the language has spread to diasporic communities through Nigerian immigrants in England, Canada, and America, amongst others. I

Multilingual Speech Recognition With A Single End-To-End Model

Training a conventional automatic speech recognition (ASR) system to support multiple languages is c

End-to-End Automatic Speech Translation of Audiobooks

We investigate end-to-end speech-to-text translation on a corpus of audiobooks specifically augmente

Large Scale Speech Recognition for Low Resource Language Amharic, an End-to-End Approach

Speech recognition, or automatic speech recognition (ASR), is a technology designed to convert spoke

Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect

Automatic speech and language technologies are still heavily biased toward high-resource languages,

End-To-End Multilingual Automatic Speech Recognition For Less-Resourced Languages: The Case Of Four Ethiopian Languages

Presenter: Solomon Teferra Abate, Martha Yifiru Tachbelie, Tanja Schultz , ICASSP 20