Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Comparative Analysis of ASR-Guided Flow-Matching and Diffusion-Based TTS for Zero-Shot Cross-Lingual Voice Cloning in

Domaine:

natural language processing

Type de record:

paper
Créateur:
Ass
Éditeur:
Zenodo
Hôte:avatar
We present PFluxTTS, a hybrid text-to-speech system addressing three gaps in flow-matching TTS: the stability-naturalness trade-off, weak cross-lingual voice cloning, and limited audio quality from low-rate mel features. Our contributions are: (1) a dual-decoder design combining duration-guided and alignment-free models through inference-time vector-field fusion; (2) robust cloning using a sequence of speech-prompt embeddings in a FLUX-based decoder, preserving speaker traits across languages without prompt transcripts; and (3) a modified PeriodWave vocoder with super-resolution to 48 kHz. On Research goal: How does the zero-shot cross-lingual voice cloning performance of ASR-guided flow-matching TTS models compare to diffusion-based architectures like Diffusion-TTS on metrics like naturalness and speaker similarity in low-resource languages? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.5/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 8.5/10.

Visit

doi.orgzenodo.org

Tasks

text to speechspeech processing

Tags

zero-shotcross-lingualvoicecloningperformanceASR-guidedflow-matchingTTS

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Comparative Analysis of Multilingual and Monolingual Intermediate-Task Training for Zero-Shot Cross-Lingual Transfer inComparative Analysis of Prefix-Tuning and Adapter-Based Fine-Tuning for Zero-Shot Cross-Lingual Generation on Low-ResourceComparative Analysis of Artificial Code-Switching and Translation Augmentation for Zero-Shot Cross-Lingual Retrieval onComparative Analysis of Optimal Transport and Contrastive Distillation for Zero-Shot Cross-Lingual Retrieval on XLSDASR data augmentation in low-resource settings using cross-lingual multi-speaker TTS and cross-lingual voice conversionComparative Analysis of Hybrid Batch Training for Zero-Shot Cross-Lingual Retrieval on MTEB Across Resource Levels

Comparative Analysis of Multilingual and Monolingual Intermediate-Task Training for Zero-Shot Cross-Lingual Transfer in

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Comparative Analysis of Prefix-Tuning and Adapter-Based Fine-Tuning for Zero-Shot Cross-Lingual Generation on Low-Resource

With the release of new large language models (LLMs) like Llama and Mistral, zero-shot cross-lingual

Comparative Analysis of Artificial Code-Switching and Translation Augmentation for Zero-Shot Cross-Lingual Retrieval on

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Comparative Analysis of Optimal Transport and Contrastive Distillation for Zero-Shot Cross-Lingual Retrieval on XLSD

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

ASR data augmentation in low-resource settings using cross-lingual multi-speaker TTS and cross-lingual voice conversion

We explore cross-lingual multi-speaker speech synthesis and cross-lingual voice conversion applied t

Comparative Analysis of Hybrid Batch Training for Zero-Shot Cross-Lingual Retrieval on MTEB Across Resource Levels

Information retrieval across different languages is an increasingly important challenge in natural l