Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Noise-Robust Multilingual Speech Recognition and the Tatar Speech Corpus

Domaine:

natural language processing

Type de record:

datasetmodel
Créateur:
SaiRinBulMan
Éditeur:
WILEY
Hôte:
After focusing on individual languages for a long time, multilingual automatic speech recognition has recently become an active area of research. For instance, Whisper by OpenAI is capable of recognizing speech in 99 languages. However, the performance of Whisper is significantly lower for lowresource languages than for high-resource ones. In this work, we aim to address this and present a fine-tuning strategy for the pretrained Whisper model so that its performance is improved for a low-resource language family while maintaining performance for a set of high-resource languages. Specifically, our Söyle model exhibited high performance for both the Turkic language family (11 languages) and the official languages of the United Nations. Our work also presents the first large open-source speech corpus for the Tatar language. We demonstrate that speech recognition performance for Tatar improves with the model trained using the new Tatar Speech Corpus (TatSC). Our model is also trained to be noise-robust. We open-source our model and TatSC to encourage further research. We envision that our fine-tuning approach will guide the creation multilingual speech recognition models for other low-resource language families.

Visit

doi.org

Tasks

automatic speech recognitionspeech processing

Licenses

https://opensource.org/licenses/MIT

Similaires

The Multilingual TEDx Corpus for Speech Recognition and TranslationA Noise-Robust End-to-End Framework for Amharic Speech RecognitionLearning Robust and Multilingual Speech RepresentationsCross-Corpus Multilingual Speech Emotion Recognition: Amharic vs. Other LanguagesThe Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech RecognitionThe Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

The Multilingual TEDx Corpus for Speech Recognition and Translation

We present the Multilingual TEDx corpus, built to support speech recognition (ASR) and speech transl

A Noise-Robust End-to-End Framework for Amharic Speech Recognition

Abstract End-to-end automatic speech recognition (ASR) offers a streamlined altern

Learning Robust and Multilingual Speech Representations

Unsupervised speech representation learning has shown remarkable success at finding representations

Cross-Corpus Multilingual Speech Emotion Recognition: Amharic vs. Other Languages

In a conventional Speech emotion recognition (SER) task, a classifier for a given language is traine

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

The Makerere AI Lab has built an end-to-end CTC Luganda ASR model using radio data. Having encountered data challenges in working with low resource languages, we take the initiative together with our partners to release the first radio corpus for Luganda. The corp

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

Building a usable radio monitoring automatic speech recognition (ASR) system is a challenging task for under-resourced languages and yet this is paramount in societies where radio is the main medium of public communication and discussions. Initial efforts by the Un