Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Low-Resource End-to-end Sanskrit TTS using Tacotron2, WaveGlow and Transfer Learning

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
DebPatNadGan
Hôte:avatar
End-to-end text-to-speech (TTS) systems have been developed for European languages like English and Spanish with state-of-the-art speech quality, prosody, and naturalness. However, development of end-to-end TTS for Indian languages is lagging behind in terms of quality. The challenges involved in such a task are: 1) scarcity of quality training data; 2) low efficiency during training and inference; 3) slow convergence in the case of large vocabulary size. In our work reported in this paper, we have investigated the use of fine-tuning the English-pretrained Tacotron2 model with limited Sanskrit data to synthesize natural sounding speech in Sanskrit in low resource settings. Our experiments show encouraging results, achieving an overall MOS of 3.38 from 37 evaluators with good Sanskrit spoken knowledge. This is really a very good result, considering the fact that the speech data we have used is of duration 2.5 hours only.

Visit

arxiv.org

Tasks

speech processingtext to speechtransfer learning

Tags

Computation and LanguageSoundAudio and Speech ProcessingMachine Learning

Similaires

Improving End-to-End Speech Translation for the Low Resource Language Fongbe to FrenchImproving noisy student training for low-resource languages in End-to-End ASR using CycleGAN and inter-domain lossesEnd-to-End Historical Handwritten Ethiopic Text Recognition Using Deep LearningLarge Scale Speech Recognition for Low Resource Language Amharic, an End-to-End ApproachAmharic OCR: An End-to-End LearningDziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect

Improving End-to-End Speech Translation for the Low Resource Language Fongbe to French

This study addresses the challenges of end-to-end (E2E) Speech-to-Text Translation (STT) for the low

Improving noisy student training for low-resource languages in End-to-End ASR using CycleGAN and inter-domain losses

Training a semi-supervised end-to-end speech recognition system using noisy student training has sig

End-to-End Historical Handwritten Ethiopic Text Recognition Using Deep Learning

Large Scale Speech Recognition for Low Resource Language Amharic, an End-to-End Approach

Speech recognition, or automatic speech recognition (ASR), is a technology designed to convert spoke

Amharic OCR: An End-to-End Learning

In this paper, we introduce an end-to-end Amharic text-line image recognition approach based on recu

Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect

Automatic speech and language technologies are still heavily biased toward high-resource languages,