Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Improving Speech Recognition for Under-resourced Languages Utilizing Audio-codecs for Data Augmentation

Domain:

natural language processing

Record type:

paper
Creator:
NirIngAnd
Publisher:
IEEE
Host:avatar
Presenter: Nirayo Hailu Gebreegziabher, Ingo Siegert, Andreas Nürnberger, MMSP 2020, Virtual Event, Sept. 21-24, 2020 To train end-to-end automatic speech recognition models, it requires a large amount of labeled speech data. This goal is challenging for languages with fewer resources. In contrast to the commonly used feature level data augmentation, we propose to expand the training set by using different audio codecs at the data level. The augmentation method consists of using different audio codecs with changed bit rate, sampling rate, and bit depth. The change reassures variation in the input data without drastically affecting the audio quality. Besides, we can ensure that humans still perceive the audio, and any feature extraction is possible later. To demonstrate the general applicability of the proposed augmentation technique, we evaluated it in an end-to-end automatic speech recognition architecture in four languages. After applying the method, on the Amharic, Dutch, Slovenian, and Turkish datasets, we achieved a 1.57 average improvement in the character error rates (CER) without integrating language models. The result is comparable to the baseline result, showing CER improvement of 2.78, 1.25, 1.21, and 1.05 for each language. On the Amharic dataset, we reached a syllable error rate reduction of 6.12 compared to the baseline result.

Visit

doi.orgrc.signalprocessingsociety.org

Tasks

automatic speech recognitionspeech processing

Languages

Amharic

Similar

Improving N-Best Rescoring in Under-Resourced Code-Switched Speech Recognition Using Pretraining and Data AugmentationAutomatic speech recognition for under-resourced languages: A surveySpeech recognition for under-resourced languages: Data sharing in hidden Markov model systemsAutomatic speech recognition for an under-resourced language - amharicText-To-Speech Data Augmentation for Low Resource Speech RecognitionData-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced Languages

Improving N-Best Rescoring in Under-Resourced Code-Switched Speech Recognition Using Pretraining and Data Augmentation

We present improvements in n-best rescoring of code-switched speech achieved by n-gram augmentation

Automatic speech recognition for under-resourced languages: A survey

(Impact-F 1.28 estim. in 2012) International audience no abstract

Speech recognition for under-resourced languages: Data sharing in hidden Markov model systems

For purposes of automated speech recognition in under-resourced environments, t

Automatic speech recognition for an under-resourced language - amharic

Text-To-Speech Data Augmentation for Low Resource Speech Recognition

Nowadays, the main problem of deep learning techniques used in the development of automatic speech r

Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced Languages

Hate speech is a global phenomenon, but most hate speech datasets so far focus on English-language c