Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Comparative Analysis of ASR-Guided Flow-Matching and Diffusion-Based TTS for Zero-Shot Cross-Lingual Voice Cloning in

Domain:

natural language processing

Record type:

paper
Creator:
Ass
Publisher:
Zenodo
Host:avatar
We present PFluxTTS, a hybrid text-to-speech system addressing three gaps in flow-matching TTS: the stability-naturalness trade-off, weak cross-lingual voice cloning, and limited audio quality from low-rate mel features. Our contributions are: (1) a dual-decoder design combining duration-guided and alignment-free models through inference-time vector-field fusion; (2) robust cloning using a sequence of speech-prompt embeddings in a FLUX-based decoder, preserving speaker traits across languages without prompt transcripts; and (3) a modified PeriodWave vocoder with super-resolution to 48 kHz. On Research goal: How does the zero-shot cross-lingual voice cloning performance of ASR-guided flow-matching TTS models compare to diffusion-based architectures like Diffusion-TTS on metrics like naturalness and speaker similarity in low-resource languages? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.5/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 8.5/10.

Visit

doi.orgzenodo.org

Tasks

text to speechspeech processing

Tags

zero-shotcross-lingualvoicecloningperformanceASR-guidedflow-matchingTTS

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Comparative Analysis of Multilingual and Monolingual Intermediate-Task Training for Zero-Shot Cross-Lingual Transfer inComparative Analysis of Prefix-Tuning and Adapter-Based Fine-Tuning for Zero-Shot Cross-Lingual Generation on Low-ResourceComparative Analysis of Artificial Code-Switching and Translation Augmentation for Zero-Shot Cross-Lingual Retrieval onComparative Analysis of Optimal Transport and Contrastive Distillation for Zero-Shot Cross-Lingual Retrieval on XLSDASR data augmentation in low-resource settings using cross-lingual multi-speaker TTS and cross-lingual voice conversionComparative Analysis of Hybrid Batch Training for Zero-Shot Cross-Lingual Retrieval on MTEB Across Resource Levels

Comparative Analysis of Multilingual and Monolingual Intermediate-Task Training for Zero-Shot Cross-Lingual Transfer in

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Comparative Analysis of Prefix-Tuning and Adapter-Based Fine-Tuning for Zero-Shot Cross-Lingual Generation on Low-Resource

With the release of new large language models (LLMs) like Llama and Mistral, zero-shot cross-lingual

Comparative Analysis of Artificial Code-Switching and Translation Augmentation for Zero-Shot Cross-Lingual Retrieval on

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Comparative Analysis of Optimal Transport and Contrastive Distillation for Zero-Shot Cross-Lingual Retrieval on XLSD

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

ASR data augmentation in low-resource settings using cross-lingual multi-speaker TTS and cross-lingual voice conversion

We explore cross-lingual multi-speaker speech synthesis and cross-lingual voice conversion applied t

Comparative Analysis of Hybrid Batch Training for Zero-Shot Cross-Lingual Retrieval on MTEB Across Resource Levels

Information retrieval across different languages is an increasingly important challenge in natural l