Logo Lanfrica

alexyigidey8/Amharic-Text-to-speech

Domaine:

natural language processing

Type de record:

model
Créateur:
ale
Hôte:
Amharic TTS with Tacotron 2 and Waveglow Amharic-Text-to-speech - Audio sample at sound cloud - sample.wav Amharic-Text-to-speech - Audio sample at sound cloud - sample.wav Main Resources Used - Tacotron 2 github.com - Waveglow phoenix.tech Training - I used a dataset of approximately 2 hours audio and text data for this results ,I used lyrics from kassmase songs as a text dataset and my voice as the audio. - **This notebook** holds the code responsibe for training the tacotron model - Directory For audio **/wavs** , Audio should be formated project rate(hz) to 22050!, you can use Audacity it's much more simpler to convert project rates(hz). - Training the model requires GPU so I recomend Colab or Kaggle The model - The accuracy of the result was pretty decent, I'd recomend providing More training data. - **This notebook** holds the code responsibe for Testing the training model - The Max-decoder variable is directly proportional to the amount of minutes the model out puts