Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

SOA: Reducing Domain Mismatch in SSL Pipeline by Speech Only Adaptation for Low Resource ASR

Domain:

natural language processing

Record type:

paper
Creator:
ShaFanAlw
Host:avatar
Recently, speech foundation models have gained popularity due to their superiority in finetuning downstream ASR tasks. However, models finetuned on certain domains, such as LibriSpeech (adult read speech), behave poorly on other domains (child or noisy speech). One solution could be collecting as much labeled and diverse data as possible for joint finetuning on various domains. However, collecting target domain speech-text paired data and retraining the model is often costly and computationally expensive. In this paper, we introduce a simple yet effective method, speech only adaptation (SOA), based on speech foundation models (Wav2vec 2.0), which requires only speech input data from the target domain. Specifically, the Wav2vec 2.0 feature encoder is continually pretrained with the Wav2vec 2.0 loss on both the source and target domain data for domain adaptation, while the contextual encoder is frozen. Compared to a source domain finetuned model with the feature encoder being frozen during training, we find that replacing the frozen feature encoder with the adapted one provides significant WER improvements to the target domain while preserving the performance of the source domain. The effectiveness of SOA is examined on various low resource or domain mismatched ASR settings, including adult-child and clean-noisy speech. Accepted to ICASSP 2024 SASB Workshop

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Languages

So_

Tags

Audio and Speech ProcessingSound

Similar

Overcoming Domain Mismatch in Low Resource Sequence-to-Sequence ASR Models using Hybrid Generated PseudotranscriptsDomain Adaptation in Low-Resource Perso-Arabic ASR with XLSR-53 Pre-TrainingDomain Robust Feature Extraction for Rapid Low Resource ASR DevelopmentDOMAIN AND LANGUAGE ADAPTATION USING HETEROGENEOUS DATASETS FOR WAV2VEC2.0-BASED SPEECH RECOGNITION OF LOW-RESOURCE LANGUAGEDomain Adaptation Effects in Cross-Lingual NER for Low-Resource LanguagesDomain Adaptation in Projection-Based NER for Low-Resource Cross-Lingual Transfer

Overcoming Domain Mismatch in Low Resource Sequence-to-Sequence ASR Models using Hybrid Generated Pseudotranscripts

Sequence-to-sequence (seq2seq) models are competitive with hybrid models for automatic speech recogn

Domain Adaptation in Low-Resource Perso-Arabic ASR with XLSR-53 Pre-Training

Self-supervised pre-training could effectively improve the performance of low-resource automatic spe

Domain Robust Feature Extraction for Rapid Low Resource ASR Development

Developing a practical speech recognizer for a low resource language is challenging, not only becaus

DOMAIN AND LANGUAGE ADAPTATION USING HETEROGENEOUS DATASETS FOR WAV2VEC2.0-BASED SPEECH RECOGNITION OF LOW-RESOURCE LANGUAGE

IEEE ICASSP 2023 Conference, Hybrid Event, 4-10 June 2023, Rhodes Island, Greece We address the effe

Domain Adaptation Effects in Cross-Lingual NER for Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Domain Adaptation in Projection-Based NER for Low-Resource Cross-Lingual Transfer

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident