Logo Lanfrica

Automatic Speech Recognition for Amharic Language using Self-Supervised

Domain:

natural language processing

Record type:

paper
Creator:
AssTam
Publisher:
Und
Host:avatar
Automatic Speech Recognition (ASR) systems have become a very natural human-machine interaction in which it allows users to speak entries rather than punching numbers on a keypad. This study presents an ASR system for Amharic language that is one of the Ethiopian Languages. The current attempts on speech recognition systems require thousands of hours of transcribed speech dataset to reach adequate accuracy. However, the large majority of under-resourced spoken languages in general, Ethiopic languages in particular, have a very limited amount of labeled speech dataset. Thus, in this study, a self-supervised Transformer based Wave2Vec 2.0 approach has been conducted to build an ASR system. In order to pretrain the proposed model in an unsupervised approach, a total of more than 200 hours unlabeled speech have been collected. In addition, a total of 30 minutes of labeled speech dataset has been prepared. Then, a Wave2Vec model that is initialized with weighted parameters has been pre-trained on the unlabeled speech dataset. Then the pretrained model has been fine-tuned using a small amount of labeled speech dataset. A standard Word Error Rate (WER) evaluation metric has been used to evaluate the fine-tuned multilingual ASR model and it has shown a clear significant result.