Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

The Effects of Input Type and Pronunciation Dictionary Usage in Transfer Learning for Low-Resource Text-to-Speech

Domaine:

natural language processing

Type de record:

paper
Créateur:
Do,ColDijKla
Hôte:avatar
We compare phone labels and articulatory features as input for cross-lingual transfer learning in text-to-speech (TTS) for low-resource languages (LRLs). Experiments with FastSpeech 2 and the LRL West Frisian show that using articulatory features outperformed using phone labels in both intelligibility and naturalness. For LRLs without pronunciation dictionaries, we propose two novel approaches: a) using a massively multilingual model to convert grapheme-to-phone (G2P) in both training and synthesizing, and b) using a universal phone recognizer to create a makeshift dictionary. Results show that the G2P approach performs largely on par with using a ground-truth dictionary and the phone recognition approach, while performing generally worse, remains a viable option for LRLs less suitable for the G2P approach. Within each approach, using articulatory features as input outperforms using phone labels. Accepted at INTERSPEECH 2023

Visit

arxiv.org

Tasks

speech processingtext to speechtransfer learning

Tags

Computation and LanguageAudio and Speech Processing

Similaires

Text-to-speech system for low-resource language using cross-lingual transfer learning and data augmentationStrategies in Transfer Learning for Low-Resource Speech Synthesis: Phone Mapping, Features Input, and Source Language SelectionAdversarial Text-to-Speech for low-resource languagesText-To-Speech Data Augmentation for Low Resource Speech RecognitionAnalyzing cross-language similarities to enhance low-resource text-to-speech via transfer learning, case study: the Moroccan Berber AmazighUsing Transfer Learning to Realize Low Resource Dungan Language Speech Synthesis

Text-to-speech system for low-resource language using cross-lingual transfer learning and data augmentation

Abstract Deep learning techniques are currently being applied in automated text-to-speech (TTS) sys

Strategies in Transfer Learning for Low-Resource Speech Synthesis: Phone Mapping, Features Input, and Source Language Selection

We compare using a PHOIBLE-based phone mapping method and using phonological features input in trans

Adversarial Text-to-Speech for low-resource languages

Improving the adversarial TTS models for low-resource languages by utilizing the high-frequency similarities between the different languages.

Text-To-Speech Data Augmentation for Low Resource Speech Recognition

Nowadays, the main problem of deep learning techniques used in the development of automatic speech r

Analyzing cross-language similarities to enhance low-resource text-to-speech via transfer learning, case study: the Moroccan Berber Amazigh

Using Transfer Learning to Realize Low Resource Dungan Language Speech Synthesis

This article presents a transfer-learning-based method to improve the synthesized speech quality of