Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Low-data? No problem: low-resource, language-agnostic conversational text-to-speech via F0-conditioned data augmentation

Domaine:

natural language processing

Type de record:

paper
Créateur:
ComHuyRibGab
Hôte:avatar
The availability of data in expressive styles across languages is limited, and recording sessions are costly and time consuming. To overcome these issues, we demonstrate how to build low-resource, neural text-to-speech (TTS) voices with only 1 hour of conversational speech, when no other conversational data are available in the same language. Assuming the availability of non-expressive speech data in that language, we propose a 3-step technology: 1) we train an F0-conditioned voice conversion (VC) model as data augmentation technique; 2) we train an F0 predictor to control the conversational flavour of the voice-converted synthetic data; 3) we train a TTS system that consumes the augmented data. We prove that our technology enables F0 controllability, is scalable across speakers and languages and is competitive in terms of naturalness over a state-of-the-art baseline model, another augmented method which does not make use of F0 information. Accepted for presentation at Interspeech 2022

Visit

arxiv.org

Tasks

text to speechspeech processing

Tags

Audio and Speech ProcessingSound

Similaires

Text-To-Speech Data Augmentation for Low Resource Speech RecognitionA Generative-Adversarial Approach to Low-Resource Language Translation via Data AugmentationText-to-speech system for low-resource language using cross-lingual transfer learning and data augmentationData Augmentation via Dependency Tree Morphing for Low-Resource Languages Text-to-Speech Synthesis Using Found Data for Low-Resource Languageskibaraki/data-augmentation-for-low-resource-asr

Text-To-Speech Data Augmentation for Low Resource Speech Recognition

Nowadays, the main problem of deep learning techniques used in the development of automatic speech r

A Generative-Adversarial Approach to Low-Resource Language Translation via Data Augmentation

Language and culture preservation is a serious challenge both socially and technologically. This pap

Text-to-speech system for low-resource language using cross-lingual transfer learning and data augmentation

Abstract Deep learning techniques are currently being applied in automated text-to-speech (TTS) sys

Data Augmentation via Dependency Tree Morphing for Low-Resource Languages

Neural NLP systems achieve high scores in the presence of sizable training dataset. Lack of such dat

Text-to-Speech Synthesis Using Found Data for Low-Resource Languages

Text-to-speech synthesis is a key component of interactive, speech-based systems. Typically, buildi

kibaraki/data-augmentation-for-low-resource-asr

Self-contained data augmentation for low-resource ASR # data-augmentation-for-low-resource-asr ##