Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

N-Shot Benchmarking of Whisper on Diverse Arabic Speech Recognition

Domain:

natural language processing

Record type:

papermodeldataset
Creator:
TalWahAbd
Host:avatar
Whisper, the recently developed multilingual weakly supervised model, is reported to perform well on multiple speech recognition benchmarks in both monolingual and multilingual settings. However, it is not clear how Whisper would fare under diverse conditions even on languages it was evaluated on such as Arabic. In this work, we address this gap by comprehensively evaluating Whisper on several varieties of Arabic speech for the ASR task. Our evaluation covers most publicly available Arabic speech data and is performed under n-shot (zero-, few-, and full) finetuning. We also investigate the robustness of Whisper under completely novel conditions, such as in dialect-accented standard Arabic and in unseen dialects for which we develop evaluation data. Our experiments show that although Whisper zero-shot outperforms fully finetuned XLS-R models on all datasets, its performance deteriorates significantly in the zero-shot setting for five unseen dialects (i.e., Algeria, Jordan, Palestine, UAE, and Yemen). 4 pages, INTERSPEECH 2023

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and LanguageSoundAudio and Speech Processing

Similar

Benchmarking Automatic Speech Recognition Models for African LanguagesAITamilDialect@DravidianLangTech 2026: Zero-Shot Whisper and Wav2Vec2 Embedding-Based Tamil Speech Recognition and Dialect Classification.Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and GermanZero-Shot Context-Aware ASR for Diverse Arabic VarietiesTamazight-Arabic Speech Recognition DatasetTamazight-Arabic Speech Recognition Dataset

Benchmarking Automatic Speech Recognition Models for African Languages

Automatic speech recognition (ASR) for African languages remains constrained by limited labeled data

AITamilDialect@DravidianLangTech 2026: Zero-Shot Whisper and Wav2Vec2 Embedding-Based Tamil Speech Recognition and Dialect Classification.

Low-resource languages pose significant challenges for speech technology due to linguistic variation

Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German

Code-switching -- the natural alternation between two languages within a single utterance -- remains

Zero-Shot Context-Aware ASR for Diverse Arabic Varieties

Zero-shot ASR for Arabic remains challenging: while multilingual models perform well on Modern Stand

Tamazight-Arabic Speech Recognition Dataset

This dataset contains speech segments in Tamazight (specifically focusing on the Tachelhit dialect)

Tamazight-Arabic Speech Recognition Dataset

This is the EMINES organization-hosted version of the Tamazight-Arabic Speech Recognition Dataset, s