Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Multilingual Training and Cross-lingual Adaptation on CTC-based Acoustic Model

Domaine:

natural language processing

Type de record:

paper
Créateur:
TonGarBou
Hôte:avatar
Multilingual models for Automatic Speech Recognition (ASR) are attractive as they have been shown to benefit from more training data, and better lend themselves to adaptation to under-resourced languages. However, initialisation from monolingual context-dependent models leads to an explosion of context-dependent states. Connectionist Temporal Classification (CTC) is a potential solution to this as it performs well with monophone labels. We investigate multilingual CTC in the context of adaptation and regularisation techniques that have been shown to be beneficial in more conventional contexts. The multilingual model is trained to model a universal International Phonetic Alphabet (IPA)-based phone set using the CTC loss function. Learning Hidden Unit Contribution (LHUC) is investigated to perform language adaptive training. In addition, dropout during cross-lingual adaptation is also studied and tested in order to mitigate the overfitting problem. Experiments show that the performance of the universal phoneme-based CTC system can be improved by applying LHUC and it is extensible to new phonemes during cross-lingual adaptation. Updating all the parameters shows consistent improvement on limited data. Applying dropout during adaptation can further improve the system and achieve competitive performance with Deep Neural Network / Hidden Markov Model (DNN/HMM) systems on limited data.

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processingtransfer learning

Tags

Audio and Speech ProcessingSound

Similaires

Comparing CTC and LFMMI for out-of-domain adaptation of wav2vec 2.0 acoustic modelOn the Analysis of Cross-Lingual Prompt Tuning for Decoder-based Multilingual ModelCode-Switching Complexity in Multilingual Model Training and Zero-Shot Cross-Lingual Retrieval PerformanceImpact of Domain Adaptation on Cross-Lingual NER Model TransferabilitySemi-supervised acoustic model training for five-lingual code-switched ASRMultilingual Intermediate-Task Training for Cross-Lingual Adversarial Robustness

Comparing CTC and LFMMI for out-of-domain adaptation of wav2vec 2.0 acoustic model

In this work, we investigate if the wav2vec 2.0 self-supervised pretraining helps mitigate the overf

On the Analysis of Cross-Lingual Prompt Tuning for Decoder-based Multilingual Model

An exciting advancement in the field of multilingual models is the emergence of autoregressive model

Code-Switching Complexity in Multilingual Model Training and Zero-Shot Cross-Lingual Retrieval Performance

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Impact of Domain Adaptation on Cross-Lingual NER Model Transferability

We propose a method for zero-resource domain adaptation of DNN acoustic models, for use in low-resou

Semi-supervised acoustic model training for five-lingual code-switched ASR

This paper presents recent progress in the acoustic modelling of under-resourced code-switched (CS)

Multilingual Intermediate-Task Training for Cross-Lingual Adversarial Robustness

Pre-trained multilingual language encoders, such as multilingual BERT and XLM-R, show great potentia