Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

XLST: Cross-lingual Self-training to Learn Multilingual Representation for Low Resource Speech Recognition

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
ZhaSonWu,Fan
Hôte:avatar
In this paper, we propose a weakly supervised multilingual representation learning framework, called cross-lingual self-training (XLST). XLST is able to utilize a small amount of annotated data from high-resource languages to improve the representation learning on multilingual un-annotated data. Specifically, XLST uses a supervised trained model to produce initial representations and another model to learn from them, by maximizing the similarity between output embeddings of these two models. Furthermore, the moving average mechanism and multi-view data augmentation are employed, which are experimentally shown to be crucial to XLST. Comprehensive experiments have been conducted on the CommonVoice corpus to evaluate the effectiveness of XLST. Results on 5 downstream low-resource ASR tasks shows that our multilingual pretrained model achieves relatively 18.6% PER reduction over the state-of-the-art self-supervised method, with leveraging additional 100 hours of annotated English data. 5 pages, 1 figure

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Audio and Speech ProcessingComputation and LanguageSound

Similaires

Cross-lingual Embedding Clustering for Hierarchical Softmax in Low-Resource Multilingual Speech RecognitionExploiting Adapters for Cross-lingual Low-resource Speech RecognitionMultilingual Intermediate-Task Training for Low-Resource Cross-Lingual TransferDonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech RecognitionLinguistic Entity Masking to Improve Cross-Lingual Representation of Multilingual Language Models for Low-Resource LanguagesMultilingual Intermediate-Task Training for Cross-Lingual Transfer in Low-Resource Languages

Cross-lingual Embedding Clustering for Hierarchical Softmax in Low-Resource Multilingual Speech Recognition

We present a novel approach centered on the decoding stage of Automatic Speech Recognition (ASR) tha

Exploiting Adapters for Cross-lingual Low-resource Speech Recognition

Cross-lingual speech adaptation aims to solve the problem of leveraging multiple rich-resource langu

Multilingual Intermediate-Task Training for Low-Resource Cross-Lingual Transfer

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition

Low-resource automatic speech recognition (ASR) commonly relies on cross-lingual transfer, where mod

Linguistic Entity Masking to Improve Cross-Lingual Representation of Multilingual Language Models for Low-Resource Languages

Multilingual Pre-trained Language models (multiPLMs), trained on the Masked Language Modelling (MLM)

Multilingual Intermediate-Task Training for Cross-Lingual Transfer in Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni