Amharic TTS with Tacotron 2 and Waveglow
Amharic-Text-to-speech
- Audio sample at sound cloud -
sample.wav
Amharic-Text-to-speech
- Audio sample at sound cloud -
sample.wav
Main Resources Used
- Tacotron 2
github.com
- Waveglow
phoenix.tech
Training
- I used a dataset of approximately 2 hours audio and text data for this results ,I used lyrics from kassmase songs as a text dataset and my voice as the audio.
- **This notebook** holds the code responsibe for training the tacotron model
- Directory For audio **/wavs** , Audio should be formated project rate(hz) to 22050!, you can use Audacity it's much more simpler to convert project rates(hz).
- Training the model requires GPU so I recomend Colab or Kaggle
The model
- The accuracy of the result was pretty decent, I'd recomend providing More training data.
- **This notebook** holds the code responsibe for Testing the training model
- The Max-decoder variable is directly proportional to the amount of minutes the model out puts