Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Fine-tuning Strategies for Faster Inference using Speech Self-Supervised Models: A Comparative Study

Domain:

natural language processing

Record type:

paper
Creator:
ZaiAlgParEss
Host:avatar
Self-supervised learning (SSL) has allowed substantial progress in Automatic Speech Recognition (ASR) performance in low-resource settings. In this context, it has been demonstrated that larger self-supervised feature extractors are crucial for achieving lower downstream ASR error rates. Thus, better performance might be sanctioned with longer inferences. This article explores different approaches that may be deployed during the fine-tuning to reduce the computations needed in the SSL encoder, leading to faster inferences. We adapt a number of existing techniques to common ASR settings and benchmark them, displaying performance drops and gains in inference times. Interestingly, we found that given enough downstream data, a simple downsampling of the input sequences outperforms the other methods with both low performance drops and high computational savings, reducing computations by 61.3% with an WER increase of only 0.81. Finally, we analyze the robustness of the comparison to changes in dataset conditions, revealing sensitivity to dataset size. Submitted to ICASSP "Self-supervision in Audio, Speech and Beyond" workshop

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Audio and Speech ProcessingMachine Learning

Similar

Fine-Tuning Strategies for Sentiment Analysis in the Algerian Dialect: A Comparative Study on DziriBERTFine-Tuned Self-Supervised Speech Representations for Language Diarization in Multilingual Code-Switched SpeechBenchmarking Self-Supervised Speech Models on Multilingual Nigerian SpeechAdapting Large Language Models to Low-Resource Tibetan: A Two-Stage Continual and Supervised Fine-Tuning StudyAutomatic Speech Recognition for Amharic Language using Self-SupervisedPretrained self-supervised speech models can recognize unseen consonants

Fine-Tuning Strategies for Sentiment Analysis in the Algerian Dialect: A Comparative Study on DziriBERT

Fine-Tuned Self-Supervised Speech Representations for Language Diarization in Multilingual Code-Switched Speech

Annotating a multilingual code-switched corpus is a painstaking process requiring specialist linguis

Benchmarking Self-Supervised Speech Models on Multilingual Nigerian Speech

Self-supervised speech models such as Whisper and wav2vec 2.0 have significantly advanced automatic

Adapting Large Language Models to Low-Resource Tibetan: A Two-Stage Continual and Supervised Fine-Tuning Study

Adapting large language models (LLMs) to low-resource languages remains a major challenge due to dat

Automatic Speech Recognition for Amharic Language using Self-Supervised

Automatic Speech Recognition (ASR) systems have become a very natural human-machine interaction in w

Pretrained self-supervised speech models can recognize unseen consonants

Modern pretrained self-supervised automatic speech recognition models are trained on large-scale aud