Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

The Usefulness of Imperfect Speech Data for ASR Development in Low-Resource Languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
Jaco BadenhorstFebe de Wet
Éditeur:
MDP
Hôte:
When the National Centre for Human Language Technology (NCHLT) Speech corpus was released, it created various opportunities for speech technology development in the 11 official, but critically under-resourced, languages of South Africa. Since then, the substantial improvements in acoustic modeling that deep architectures achieved for well-resourced languages ushered in a new data requirement: their development requires hundreds of hours of speech. A suitable strategy for the enlargement of speech resources for the South African languages is therefore required. The first possibility was to look for data that has already been collected but has not been included in an existing corpus. Additional data was collected during the NCHLT project that was not included in the official corpus: it only contains a curated, but limited subset of the data. In this paper, we first analyze the additional resources that could be harvested from the auxiliary NCHLT data. We also measure the effect of this data on acoustic modeling. The analysis incorporates recent factorized time-delay neural networks (TDNN-F). These models significantly reduce phone error rates for all languages. In addition, data augmentation and cross-corpus validation experiments for a number of the datasets illustrate the utility of the auxiliary NCHLT data.

Visit

doi.org

Tasks

automatic speech recognitionspeech processing

Licenses

https://creativecommons.org/licenses/by/4.0/

Similaires

elerdg/ASR-for-low-resource-languagesAutomatic Speech Recognition (ASR) for African Low-Resource Languages: A Systematic Literature ReviewTowards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource LanguagesEfficient ASR for Low-Resource Languages: Leveraging Cross-Lingual Unlabeled DataMultimodal In-context Learning for ASR of Low-resource Languages Text-to-Speech Synthesis Using Found Data for Low-Resource Languages

elerdg/ASR-for-low-resource-languages

Fine-tune wav2vec2-xls-r on data from low-resource-languages # ASR for Low-resource languages ## O

Automatic Speech Recognition (ASR) for African Low-Resource Languages: A Systematic Literature Review

ASR has achieved remarkable global progress, yet African low-resource languages remain rigorously un

Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages

Automatic speech recognition (ASR) for dysarthric speech remains challenging due to data scarcity, p

Efficient ASR for Low-Resource Languages: Leveraging Cross-Lingual Unlabeled Data

Automatic speech recognition for low-resource languages remains fundamentally constrained by the sca

Multimodal In-context Learning for ASR of Low-resource Languages

Automatic speech recognition (ASR) still covers only a small fraction of the world's languages, main

Text-to-Speech Synthesis Using Found Data for Low-Resource Languages

Text-to-speech synthesis is a key component of interactive, speech-based systems. Typically, buildi