Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

How Does Pre-trained Wav2Vec 2.0 Perform on Domain Shifted ASR? An Extensive Benchmark on Air Traffic Control Communications

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
ZulPraNigSar
Hôte:avatar
Recent work on self-supervised pre-training focus on leveraging large-scale unlabeled speech data to build robust end-to-end (E2E) acoustic models (AM) that can be later fine-tuned on downstream tasks e.g., automatic speech recognition (ASR). Yet, few works investigated the impact on performance when the data properties substantially differ between the pre-training and fine-tuning phases, termed domain shift. We target this scenario by analyzing the robustness of Wav2Vec 2.0 and XLS-R models on downstream ASR for a completely unseen domain, air traffic control (ATC) communications. We benchmark these two models on several open-source and challenging ATC databases with signal-to-noise ratio between 5 and 20 dB. Relative word error rate (WER) reductions between 20% to 40% are obtained in comparison to hybrid-based ASR baselines by only fine-tuning E2E acoustic models with a smaller fraction of labeled data. We analyze WERs on the low-resource scenario and gender bias carried by one ATC dataset. To be published in the 2022 IEEE Spoken Language Technology Workshop (SLT) (SLT 2022)

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Audio and Speech ProcessingComputation and LanguageMachine Learning

Similaires

Nicolingua Pre-Trained Model: West African Wav2vecStrategies for improving low resource speech to text translation relying on pre-trained ASR modelsComparing CTC and LFMMI for out-of-domain adaptation of wav2vec 2.0 acoustic modelDeepfake Speech on African Accents: How do Modern Systems Perform?How Linguistically Fair Are Multilingual Pre-Trained Language Models?Zero-shot cross-lingual transfer of domain-diverse pre-trained models on XTREME-R for low-resource versus high-resource languages

Nicolingua Pre-Trained Model: West African Wav2vec

Trained on the West African Radio Corpus.

Strategies for improving low resource speech to text translation relying on pre-trained ASR models

This paper presents techniques and findings for improving the performance of low-resource speech to

Comparing CTC and LFMMI for out-of-domain adaptation of wav2vec 2.0 acoustic model

In this work, we investigate if the wav2vec 2.0 self-supervised pretraining helps mitigate the overf

Deepfake Speech on African Accents: How do Modern Systems Perform?

Deepfake Speech on African Accents: How do Modern Systems Perform?

Poster presented at the Deep Learning Indaba 2023 by Kweku Andoh Yamoah

How Linguistically Fair Are Multilingual Pre-Trained Language Models?

Massively multilingual pre-trained language models, such as mBERT and XLM-RoBERTa, have received sig

Zero-shot cross-lingual transfer of domain-diverse pre-trained models on XTREME-R for low-resource versus high-resource languages

Pre-trained multilingual language encoders, such as multilingual BERT and XLM-R, show great potentia