Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

ASR data augmentation in low-resource settings using cross-lingual multi-speaker TTS and cross-lingual voice conversion

Domaine:

natural language processing

Type de record:

paper
Créateur:
CasShuKorJun
Hôte:avatar
We explore cross-lingual multi-speaker speech synthesis and cross-lingual voice conversion applied to data augmentation for automatic speech recognition (ASR) systems in low/medium-resource scenarios. Through extensive experiments, we show that our approach permits the application of speech synthesis and voice conversion to improve ASR systems using only one target-language speaker during model training. We also managed to close the gap between ASR models trained with synthesized versus human speech compared to other works that use many speakers. Finally, we show that it is possible to obtain promising ASR training results with our data augmentation method using only a single real speaker in a target language. This paper was accepted at INTERSPEECH 2023

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Audio and Speech ProcessingComputation and LanguageSound

Similaires

Text-to-speech system for low-resource language using cross-lingual transfer learning and data augmentationEfficient ASR for Low-Resource Languages: Leveraging Cross-Lingual Unlabeled DataQuery Augmentation for Cross-Lingual Dense Retrieval in Low-Resource LanguagesCross-lingual NER Performance with Unlabeled Target Data in Low-Resource SettingsVoice Conversion Can Improve ASR in Very Low-Resource SettingsMulti-Positive Contrastive Learning with Synthetic Misspelling Augmentation for Cross-Lingual Dense Retrievers in Low-Resource

Text-to-speech system for low-resource language using cross-lingual transfer learning and data augmentation

Abstract Deep learning techniques are currently being applied in automated text-to-speech (TTS) sys

Efficient ASR for Low-Resource Languages: Leveraging Cross-Lingual Unlabeled Data

Automatic speech recognition for low-resource languages remains fundamentally constrained by the sca

Query Augmentation for Cross-Lingual Dense Retrieval in Low-Resource Languages

Effective cross-lingual dense retrieval methods that rely on multilingual pre-trained language model

Cross-lingual NER Performance with Unlabeled Target Data in Low-Resource Settings

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Voice Conversion Can Improve ASR in Very Low-Resource Settings

Voice conversion (VC) could be used to improve speech recognition systems in low-resource languages

Multi-Positive Contrastive Learning with Synthetic Misspelling Augmentation for Cross-Lingual Dense Retrievers in Low-Resource

Dense retrieval has become the new paradigm in passage retrieval. Despite its effectiveness on typo-