Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text

Domaine:

natural language processing

Type de record:

papermodelsoftware
Créateur:
Li,Pu,SunZha
Hôte:avatar
Whisper and other large-scale automatic speech recognition models have made significant progress in performance. However, their performance on many low-resource languages, such as Kazakh, is not satisfactory. It is worth researching how to utilize low-cost data to improve the performance of Whisper on under-represented languages. In this study, we utilized easily accessible unpaired speech and text data and combined the language model GPT with Whisper on Kazakh. We implemented end of transcript (EOT) judgment modification and hallucination penalty to improve the performance of speech recognition. Further, we employed the decoding average token log probability as a criterion to select samples from unlabeled speech data and used pseudo-labeled data to fine-tune the model to further improve its performance. Ultimately, we achieved more than 10\% absolute WER reduction in multiple experiments, and the whole process has the potential to be generalized to other under-represented languages. Accepted by INTERSPEECH 2024;Minor typo correction

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Audio and Speech ProcessingComputation and LanguageSound

Similaires

Speech Language Models for Under-Represented Languages: Insights from WolofUsing Songs to Improve Kazakh Automatic Speech RecognitionDoro97/African-language-Speech-Recognition---Speech-to-Text-smegnshd/African-language-Speech-Recognition---Speech-to-Text-Rukundo725/African-language-Speech-Recognition---Speech-to-TextSpeech Processing for Text Independent Amharic Language Dialect Recognition

Speech Language Models for Under-Represented Languages: Insights from Wolof

We present our journey in training a speech language model for Wolof, an underreprese

Using Songs to Improve Kazakh Automatic Speech Recognition

Developing automatic speech recognition (ASR) systems for low-resource languages is hindered by the

Doro97/African-language-Speech-Recognition---Speech-to-Text-

# African-language-Speech-Recognition---Speech-to-Text- The World Food Program wants to deploy an i

smegnshd/African-language-Speech-Recognition---Speech-to-Text-

# Introduction The World Food Program wants to deploy an intelligent form that collects nutritional

Rukundo725/African-language-Speech-Recognition---Speech-to-Text

# African-language-Speech-Recognition---Speech-to-Text The design of this intelligent form require

Speech Processing for Text Independent Amharic Language Dialect Recognition