Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

DRAFT: A Novel Framework to Reduce Domain Shifting in Self-supervised Learning and Its Application to Children's ASR

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
FanAlw
Hôte:avatar
Self-supervised learning (SSL) in the pretraining stage using un-annotated speech data has been successful in low-resource automatic speech recognition (ASR) tasks. However, models trained through SSL are biased to the pretraining data which is usually different from the data used in finetuning tasks, causing a domain shifting problem, and thus resulting in limited knowledge transfer. We propose a novel framework, domain responsible adaptation and finetuning (DRAFT), to reduce domain shifting in pretrained speech models through an additional adaptation stage. In DRAFT, residual adapters (RAs) are inserted in the pretrained model to learn domain-related information with the same SSL loss as the pretraining stage. Only RA parameters are updated during the adaptation stage. DRAFT is agnostic to the type of SSL method used and is evaluated with three widely used approaches: APC, Wav2vec2.0, and HuBERT. On two child ASR tasks (OGI and MyST databases), using SSL models trained with un-annotated adult speech data (Librispeech), relative WER improvements of up to 19.7% are observed when compared to the pretrained models without adaptation. Additional experiments examined the potential of cross knowledge transfer between the two datasets and the results are promising, showing a broader usage of the proposed DRAFT framework. Accepted to Interspeech 2022

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Audio and Speech ProcessingSound

Similaires

End-to-end Jordanian dialect speech-to-text self-supervised learning frameworkSemi-Supervised Learning to Perceive Children's Affective States in a Tablet TutorLearning to allocate: Self-supervised transformers for constrained optimisationSelf-supervised and Multilingual Learning Applied to the Wolof, Swahili and FongbeFast Development of ASR in African Languages using Self Supervised Speech Representation LearningFrom speech to primate vocalizations: self-supervised deep learning as a comparative approach

End-to-end Jordanian dialect speech-to-text self-supervised learning framework

Speech-to-text engines are extremely needed nowadays for different applications, representing an ess

Semi-Supervised Learning to Perceive Children's Affective States in a Tablet Tutor

Like good human tutors, intelligent tutoring systems should detect and respond to students' affectiv

Learning to allocate: Self-supervised transformers for constrained optimisation

We present the Resource Allocation Transformer, a deep learning framework that learns portfolio-leve

Self-supervised and Multilingual Learning Applied to the Wolof, Swahili and Fongbe

Fast Development of ASR in African Languages using Self Supervised Speech Representation Learning

This paper describes the results of an informal collaboration launched during the African Master of

From speech to primate vocalizations: self-supervised deep learning as a comparative approach

International audience The deep learning revolution partly embodied in transformers a