Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Distilling a Pretrained Language Model to a Multilingual ASR Model

Domain:

natural language processing

Record type:

papermodel
Creator:
ChoPar
Host:avatar
Multilingual speech data often suffer from long-tailed language distribution, resulting in performance degradation. However, multilingual text data is much easier to obtain, yielding a more useful general language model. Hence, we are motivated to distill the rich knowledge embedded inside a well-trained teacher text model to the student speech model. We propose a novel method called the Distilling a Language model to a Speech model (Distill-L2S), which aligns the latent representations of two different modalities. The subtle differences are handled by the shrinking mechanism, nearest-neighbor interpolation, and a learnable linear projection layer. We demonstrate the effectiveness of our distillation method by applying it to the multilingual automatic speech recognition (ASR) task. We distill the transformer-based cross-lingual language model (InfoXLM) while fine-tuning the large-scale multilingual ASR model (XLSR-wav2vec 2.0) for each language. We show the superiority of our method on 20 low-resource languages of the CommonVoice dataset with less than 100 hours of speech data. Accepted to Interspeech 2022. Official implementation provided in github.com

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and LanguageArtificial IntelligenceSoundAudio and Speech Processing

Similar

Learning ASR pathways: A sparse multilingual ASR modelAfroLM: A Self-Active Learning-based Multilingual Pretrained Language Model for 23 African LanguagesDN at SemEval-2023 Task 12: Low-Resource Language Text Classification via Multilingual Pretrained Language Model Fine-tuningyemen2016/ewondo-pretrained-modelFinetuning Pretrained Model with Embedding of Domain and Language Information for ASR of Very Low-Resource SettingsEwe Language ASR Dataset & Model

Learning ASR pathways: A sparse multilingual ASR model

Neural network pruning compresses automatic speech recognition (ASR) models effectively. However, in

AfroLM: A Self-Active Learning-based Multilingual Pretrained Language Model for 23 African Languages

In recent years, multilingual pre-trained language models have gained prominence due to their remark

DN at SemEval-2023 Task 12: Low-Resource Language Text Classification via Multilingual Pretrained Language Model Fine-tuning

In recent years, sentiment analysis has gained significant importance in natural language processing

yemen2016/ewondo-pretrained-model

Finetuning Pretrained Model with Embedding of Domain and Language Information for ASR of Very Low-Resource Settings

This study investigates the effective incorporation of meta-information such as domain and language

Ewe Language ASR Dataset & Model

Voice-based agricultural extension, digital health tools, AI chatbots in local languages, language m